← ← Back to Supply Chain Review Logistics Technology

Reply sets 5 authority levels for warehouse AI agents; five go live on 5 October

Source: IT Tech News · 2026-10-07
中文
Summary

Logistics Reply launched the LEA AI Agent Authority Model on 5 October, crossing four stages of org AI maturity with five authority levels - Inform, Recommend, Act, Coordinate and Governed Autonomy. The premise is blunt: able to act is not reason to let it. Five pre-built agents shipped the same day, no custom dev: Out of Stock, Labour Distribution, ABC Rebalancer, Dock Scheduling and Lost & Found, which reads camera input to find what blocked an AMR.

Supply Chain Action Points

Logistics Reply's LEA AI Agent Authority Model, launched 5 October, is less a product and more a way of thinking about who gets to do what when you put software agents inside a warehouse. The headline idea is uncomfortable for anyone who has been sold the full-automation dream: just because an agent can act, that does not mean it should be allowed to. Reply ships the model with five ready-made agents and five authority levels, and the interesting move is that you get to decide, per task, how far each agent is trusted. For importers, exporters and the 3PL warehouses that run their stock, the practical question is not whether to buy this, but where to start and what to keep a human hand on.

**What the five-level model actually hands a buyer**

Most warehouse AI pitches show you the autonomous end state and leave the messy middle undefined. Reply's Authority Model does the opposite: it puts the middle steps — the part where you decide how much rope to give the software — front and center. You get five authority levels, and the explicit message is that you choose where each agent sits, not the vendor. For an importer or exporter paying a 3PL to run stock, that choice is the actual product. You are not buying a robot that does whatever it decides; you are buying a set of agents whose leash length you set, lane by lane, task by task.

The five levels are Inform, Recommend, Act, Coordinate and Governed Autonomy. Inform is the bottom rung: the agent watches, notices, and tells you, but the decision stays with a human. A good example is an agent that flags a bin approaching its reorder point and leaves the purchase order to you. Recommend goes one step up: the agent now hands you a proposal with its reasoning and waits for a click. Act is where the agent executes, but inside a fence you drew — fixed rules, fixed limits, a defined set of things it is allowed to touch, and usually a defined set it is forbidden from touching. Coordinate means the agent is no longer working one task in isolation; it is sequencing several agents or steps so they do not collide over the same resource. Governed Autonomy is the top: the agent runs on its own, within governance someone defined up front and that still gets reviewed.

The relief, if you run a warehouse, is that nobody is forcing the autonomous demo down your throat. The danger worth naming out loud is the opposite one: vendors love the Governed Autonomy headline, and procurement teams sometimes sign for it without asking who eats the cost when an agent reroutes a shipment wrong at 3 a.m. The value of Reply's model is that it gives you the language to say not yet and mean it, and to write that not yet into a record someone checks.

**Why able to act is not the same as should let it act**

The premise Reply leads with is blunt and worth repeating: able to act is not a reason to let it. A system can be technically capable of auto-releasing a dock appointment, but that does not make auto-release the right call on a site where carrier no-shows run high and a wrong slot backs up the whole yard for the day. Capability and appropriateness are two different questions, and the model is built to keep them separate instead of letting the demo answer both.

Think about it from the floor. A labour-distribution agent that can shuffle shifts sounds efficient until it moves a certified forklift operator off a task that needs the certification, or splits a pick team mid-run and drops their rate for an hour while they re-form. The agent did exactly what it was allowed to do. The human paid for it in missed cutoffs and a grumpy crew. The fix is not a smarter agent; it is a tighter fence around when the agent is allowed to act at all, and on which SKUs, and during which shift.

There is also a liability shape most pilots ignore until it bites. When an agent only recommends, the human owns the decision and the audit trail is clean: here is what it said, here is what we did, here is who signed. The moment the agent acts on its own, the responsibility diagram gets muddy, and your insurance language, your SOPs, and your incident process all need to catch up before the technology does. Reply shipping five pre-built agents on day one does not mean five agents should be at Act on day one. It means five chances to practice saying no, and to say it with a reason attached.

**How to stage a rollout without betting the warehouse**

The five levels double as a rollout sequence, and you should treat them that way rather than jumping to the top because the brochure is prettiest there. Start at Inform for everything. Let the agents watch for a week or two and report what they see. You are calibrating your own trust and collecting evidence about whether their read of the warehouse matches the reality on the floor. Resist the urge to promote anything just because it is accurate at reporting — reporting accuracy is the easy bar.

Move to Recommend next. Now the agent proposes and a human approves. This is where you learn whether its reasoning holds up under someone who actually knows the site, not under a benchmark. If it keeps recommending things a shift lead would never do, that is data, not failure — it tells you the fence is wrong, the data feed is thin, or the agent is solving a problem the floor does not have. All three are fixable, and all three are cheaper to find at Recommend than at Act.

Promote to Act only for narrow, reversible, low-blast-radius tasks. Reordering a cheap fast-mover inside a fixed band is a good Act candidate: the wrong move is a small overstock a human catches and cancels. Auto-cancelling a carrier booking is not a good candidate, because the blast radius reaches outside your four walls and into a customer's delivery promise. Coordinate comes later, once individual agents have earned Act on their own lanes and you are confident they will not fight each other for the same dock door or the same operator. Governed Autonomy is the last step, and for most mid-size 3PLs it may stay a someday goal rather than a this-quarter target. There is no prize for reaching it fast, and a wrong autonomous decision at scale is the most expensive mistake on the list.

**Which of the five agents to pilot first, and why**

Reply shipped five pre-built agents with no custom development: Out of Stock, Labour Distribution, ABC Rebalancer, Dock Scheduling, and Lost & Found, the last of which reads camera input to find what blocked an AMR. They are not equal in risk, so the pilot order should not be alphabetical or based on which demo looked slickest.

Out of Stock is the easiest win. It watches inventory positions and flags or acts on gaps. The worst case of a wrong move is a premature replenishment signal, which a human can catch and cancel before money moves. That makes it a safe Recommend-then-Act candidate and a good first pilot for almost any site.

ABC Rebalancer rebalances stock by ABC classification, putting fast movers where they are cheapest to pick. Wrong moves are annoying but reversible, and the payoff in pick-time savings shows up quickly on a busy site. Pilot it second, right behind Out of Stock.

Lost & Found is the interesting one. It reads camera input to find what blocked an AMR, which is genuinely useful on robot-heavy floors where a stray pallet or a misplaced cart can halt a whole aisle. Start it at Inform. Let it tell you a pallet at aisle 7 blocked the robot, verify it is right a few dozen times against what really happened, and only then consider promoting it. Its value is real, but trusting camera interpretation too early is how you get a robot told to route around a shadow or a forklift that was supposed to be there.

Labour Distribution should wait. Moving people is the highest-blame, hardest-to-undo action in the set, and it touches contracts, certifications, and morale. Keep it on Recommend for a long time, and only let it act inside rules a supervisor already wrote and signs off on weekly.

Dock Scheduling is the one I would keep on human-in-the-loop essentially permanently, or at minimum require a dispatcher's sign-off on every move. It sits at the junction of carriers, yard space, and delivery SLAs, and a bad auto-decision ripples outward across the whole day's plan. Let the agent recommend the slot; let a person confirm it. If you ever promote it, promote it last and only on a site with clean, system-of-record slot data.

**Governance guardrails you need before any agent reaches Act**

Before you flip any agent from Recommend to Act, write down three things in one place: what it is allowed to do, what it is forbidden from doing, and what happens when it is wrong. The third one is the one everyone skips, and it is the only one that matters at 2 a.m. when a replenishment fired and nobody is awake to catch it. What happens when it is wrong needs a named owner and a named recovery step, not a vague hope.

You also need an authority register — a single page, or a single screen, listing every agent, its current level, who approved that level, and when it gets reviewed. Without it, agents drift upward as people get comfortable, and six months later nobody can tell you which ones are autonomous and which are still asking permission. Comfort is not a governance process, and comfort is exactly what quietly promotes an agent.

And keep a human override that is one click, not a ticket. If a shift lead cannot switch an agent back to Recommend in seconds, your guardrail is theoretical. The whole point of the five-level model is that authority is a dial, and a dial you cannot turn is just a locked setting wearing a friendly label. Test the override weekly. If it takes longer than finding a supervisor, it does not count.

**A 30-day playbook**

Week one: stand up all five agents at Inform. Read their reports. Do not let any of them act. Your only job is to learn whether their view of the warehouse matches yours, and to note where they disagree with the floor's intuition.

Week two: promote Out of Stock and ABC Rebalancer to Recommend. A human approves every move. Track how often the human agrees, and write down the cases where they did not — those disagreements are your fence-tuning inputs, not noise.

Week three: if the agreement rate is high and the wrong-call cost is low, let Out of Stock act within a tight band — cheap, fast, reversible items only, with a hard ceiling. Keep ABC Rebalancer on Recommend a little longer. Start Lost & Found at Inform and begin verifying its camera reads against reality, a dozen checks a day.

Week four: review. Write the authority register. Decide which agent, if any, earns Act next, and hold Labour Distribution and Dock Scheduling at Recommend. Confirm the one-click human override works on every live agent, including the ones still at Inform, because you will need it the moment something looks wrong.

**Worked example: a mid-size 3PL, mapped out**

Assume a 3PL with roughly 18,000 square meters of ambient storage, about 4,000 active SKUs, two shifts, and a modest AMR fleet used for putaway. Assume its data is decent: the WMS keeps clean bin and transaction records, but labour scheduling lives mostly in a supervisor's head, and dock slots are booked by email rather than in a system of record. Those last two are the weak spots, and they shape the whole rollout, because an agent can only be as good as the data behind it.

Under the model, this site would launch all five agents at Inform in week one and simply read. By week two, Out of Stock and ABC Rebalancer move to Recommend; the supervisor approves and the system logs every yes and no. Because the labour data is informal, Labour Distribution stays at Inform longer than the others — you cannot safely recommend shift moves from a head that is not yet a feed. Dock Scheduling also stays low, because email-booked slots are a messy source of truth, and an agent recommending into that mess would mostly be guessing with extra steps.

By week four, Out of Stock earns Act within a fixed reorder band for C-class slow movers only — low value, low risk, easy to reverse. ABC Rebalancer stays at Recommend because rebalancing A-class fast movers wrong is expensive. Lost & Found runs at Inform, reporting blockages, while the team checks its camera calls against reality and tunes the confidence threshold. Labour Distribution and Dock Scheduling remain on human-in-the-loop. The assumption baked in, and stated out loud to the client, is that this site is not ready for Coordinate or Governed Autonomy this quarter, and forcing it would trade a small efficiency gain for a real operational risk the client would end up owning.

**What to monitor once agents are live**

Watch four numbers, not forty. Start with the agreement rate between human and agent at the Recommend level — if it drops, the agent's model or your data has drifted, and drift is silent until it is expensive. Next, time-to-catch for any wrong Act-level move: how fast did a human notice and reverse it, and was the reversal before or after the cost hit. Then, exception volume — a sudden spike in agent flags usually means a feed broke, not that the warehouse exploded, so treat a spike as a plumbing alarm first. After that, override frequency — if supervisors keep switching agents back to Recommend, the level is set too high for this site, and that is a signal to pull the dial back, not a reason to lecture the floor.

None of this needs a dashboard the size of a wall. A weekly five-minute read of those four numbers tells you whether an agent is ready to climb a level or needs to come back down. The model gives you the dials; monitoring is how you know which dial to turn, and when, and by how much.

**What this means for your contract with the 3PL**

If you are an importer or exporter, the agent question lands on you indirectly. Your 3PL might adopt LEA-style authority models, and suddenly who decided this replenishment has a new answer that is not a person. You should ask, in plain terms, which agents touch your stock and at what level. A 3PL running Out of Stock at Act on your fast movers is a different risk than one running it at Recommend, and you are the one who eats a stockout or an overstock either way. Put the authority level in the contract or an SOP appendix, even informally, because it is cheaper to agree on liability before an agent misroutes a container than after.

There is also a service-level angle. If your 3PL promotes Dock Scheduling to autonomous too early and starts stacking carriers badly, your downstream deliveries slip and your customer blames you, not the software. So the cautious 3PL is the better partner, and the five-level model gives you a vocabulary to reward caution instead of penalizing it. When a provider tells you we keep labour and dock on human-in-the-loop, that is a maturity signal, not a weakness to be negotiated away.

**Data readiness: the quiet prerequisite**

None of the five levels work if the feeds are thin, and this is the part that gets understated in every sales conversation. An Out of Stock agent is only as good as the transaction and on-hand data feeding it; a Labour Distribution agent is only as good as the labour and certification records behind it; a Dock Scheduling agent is only as good as the slot data it reads. Before you promote anything past Inform, spend real time on the inputs. Dirty bin records produce confident, wrong recommendations, and the agent will not know it is wrong, because confidence and correctness are not the same thing.

This is exactly why the worked example above held Labour Distribution back: the supervisor's head is not a feed. Until that data is captured in a system the agent can read, recommending shift moves is guessing with extra steps and a friendly interface. The lesson generalizes across all five agents — audit the input before you trust the output, at every level, and re-audit when a feed changes owner or format.

**Change management for the floor**

Agents change who does what, and the floor feels that first, before the dashboard does. The honest version, said out loud by management: an agent at Recommend does not replace the shift lead, it changes what the lead spends time on. Instead of building the replenishment list by hand, the lead reviews one someone else drafted. That is a better job, not a smaller one, but people need to hear it from a manager, not infer it from a screen that suddenly has opinions.

Bring the leads into the pilot decisions. When a supervisor helps draw the fence around an agent, they trust the agent more and catch its errors faster, because they know exactly where the line was drawn and why. The five-level model is, in practice, a standing negotiation between the software's capability and the floor's judgment, and the floor should have a seat at that table from week one, not after something goes wrong.

**The cost of being wrong, by agent**

A useful exercise before promoting any agent is to write the worst plausible wrong move for each, and what it costs to undo. Out of Stock fires a false reorder on a cheap item: a small overstock, easy to fix, low stakes. ABC Rebalancer puts a fast mover in a far bin: extra pick seconds across every cycle, recoverable but annoying at volume. Lost & Found misreads the camera and sends an AMR the wrong way: a stoppage, and on a busy floor a collision risk you do not want to find out about. Labour Distribution moves a certified operator off a certified task: a compliance gap and a rework, the kind of mistake that shows up in an audit rather than a daily report. Dock Scheduling books a slot the carrier cannot make: a yard backup and a missed cutoff that ripples to customers. The list ranks the promotion order by itself — cheapest to undo first, most expensive last.

**Don't confuse the model with the agents**

One more distinction worth keeping straight: the Authority Model is the framework, and the five agents are the first occupants. The model is the part that outlasts any single agent, any single vendor release, any single site. A buyer should evaluate them separately. You might love the model and still find that your operation does not need Lost & Found because you have no AMRs, or that you want the model's discipline applied to agents you build yourself later. The launch bundles them, but they are separable ideas, and treating them as one is how you over-buy a capability you then under-use.

The practical takeaway for an importer, exporter, or 3PL operator is narrower than the press release suggests. You do not need to adopt all five agents, and you do not need to reach Governed Autonomy to get value. You need the habit of setting authority deliberately, per agent, per lane, and revisiting it on a schedule. That habit — more than any one agent — is what keeps able to act from quietly becoming should have let it.

  • Start every agent at Recommend, never at Act, for the first 30 days; keep the human in the yes/no loop and log each decision so you build an evidence trail.
  • Gate the Act level behind a bar you define yourself: only promote an agent after it hits your agreed accuracy over a set number of Recommend decisions, and only on reversible, low-blast-radius tasks.
  • Keep Dock Scheduling on human-in-the-loop permanently, or at minimum require a dispatcher's sign-off on every move, because it sits at the carrier, yard and SLA crossroads.
  • Run Lost & Found at Inform first and verify its camera reads against reality a few dozen times before trusting it to do anything beyond reporting.
  • Write a one-page authority register listing every agent, its current level, who approved it, and the review date; review it monthly, not when something breaks.
  • Pilot Out of Stock and ABC Rebalancer before Labour Distribution, because a wrong inventory move is far cheaper to undo than a wrong labour move.

— 作者 Ravi Chandran

Read original article →
agentic-aiwarehousegovernanceautomation