← Back to Supply Chain Review Logistics Technology

Qingdao Port's Star Harbor Model Cuts Berth Planning from Hours to 90 Seconds

Source: EEO · 2026-10-11 · 16 min read
中文

Supply Chain Action Points

Read this first — the conclusion, and the moves to make:

  1. Pull 12 months of forecast ETA vs actual berthing time this month, bucketed at 72/48/24/12 hours out, with MAE and sigma per bucket — no optimisation spend before this baseline exists.
  2. Write the data calibre down by end of quarter: who enters the ETA, who updates it, who owns the discharge-rate estimate, and which one is the single source of truth when they disagree.
  3. Hold inputs to explicit thresholds: ETA MAE at or under 3 hours at 48 hours out, at or under 1 hour at 12 hours out, discharge-rate estimate within 10% of actual, dynamic fields refreshed every 15 minutes.
  4. Replace the adoption metric within two quarters: move from 'taken more than 90% of the time' to zero-modification adoption at 70% or better, with the manual edit log instrumented from this month.
  5. Put a churn guardrail on the re-solve now: human confirmation whenever a new plan shifts a vessel by more than one tidal window or swaps more than two berth assignments.
  6. Ask your agent for an arrival distribution rather than a date, feed that sigma into your safety stock formula, and put P95 waiting time — not the average — into the laytime clause at the next contract renewal.
Skip to the detailed analysis ↓
Summary

Qingdao's Star Harbor model, released on 31 October 2025 as China's first port-industry large language model, now runs 26 scenarios and 19 agents. On dry bulk, its scheduling agent weighs 132 elements against more than 180 business rules and returns an optimal berthing plan in 90 seconds, where a veteran planner needed two to three hours and eight to ten years of training; recommended plans are taken over 90% of the time. Qingdao reports capacity up 10%, yard re-handling down 5% and energy cost down 8%.

The Analysis

Every port story that ships with a stopwatch in the headline deserves a second read, and this one ships with a very loud stopwatch. Qingdao's Star Harbor model went live on 31 October 2025 as the first port-industry large language model in China, and it now runs across 26 demonstration scenarios and 19 agents. The piece everybody is quoting is the dry bulk scheduling agent: it weighs 132 planning elements against more than 180 business rules and returns an optimal berthing plan in 90 seconds, where a veteran planner needed two to three hours. Training that planner to maturity took eight to ten years. The recommended plan is taken more than 90% of the time. Qingdao also puts berth throughput up 10%, yard re-handling down 5% and energy cost down 8%.

I do not read this as a speed story. I read it as a queue.

Let me model it before anybody argues about it, because the interesting part is not the 90 seconds. A berth plan is a queueing system with a scheduling layer bolted on top: vessels arrive on a stochastic arrival process, berths and cranes are the servers, service time is the discharge rate in tonnes per hour, and the tide-plus-channel window behaves like a gate that only opens in batches. Anyone who has sat at anchor outside a bulk terminal has felt this system in their bones. The question that matters is not who computes faster. The question is which parameter moves the whole thing once you touch it.

Start with what actually changed, in the driest possible terms. Two of the numbers in that release are hard and one of them is soft. Hard: 132 planning elements and more than 180 business rules now get weighed together, and the answer comes back in 90 seconds. Also hard: the plan is taken more than 90% of the time. Soft: the two to three hours it used to take. I call that last one soft because it is an average taken across good days and bad days, across a planner who has had coffee and a planner who has been awake since two in the morning, and because nobody ever publishes the variance behind it. Someone who thinks in queues hears two to three hours and immediately wants the distribution, not the mean.

Here is the model I would build. Arrivals: vessels show up at the pilot station carrying an ETA that has error in it, and that error grows the further out you look. Servers: berths, and behind the berths the cranes, and behind the cranes the yard, and behind the yard the rail slots and the truck gates. Service time: the discharge rate, which is not a constant but a function of hatch count, grab condition, weather, and how badly the stow plan was written five ports ago. Constraint: the tidal and channel window, which effectively batches arrivals into waves. Your plan is an assignment problem — which vessel takes which berth in which window, given that moving one vessel moves everyone queued behind it.

So what did the 90 seconds actually buy? Not throughput on its own. Ninety seconds solved a problem that was never the binding constraint. A berth plan gets built a handful of times a day and revised when something breaks. Two hours of compute on a task that runs maybe six times a day is simply not where the money goes. If the port had bought speed alone, the honest return would be a few planner-hours a day and the whole investment would be hard to defend.

Let me put a number on it, because it is the number nobody prints. Assume four planners, six plans a day each, 2.5 hours saved per plan by going from roughly two and a half hours to near zero. That is about 57.6 planner-hours a day, call it 21,000 hours a year. At a fully loaded 30 dollars an hour — my assumption, not their figure — that is roughly 630,000 dollars a year. Real money, small money. It does not pay for a large language model.

Where the value actually sits is two places the release does not headline. One is consistency. A model produces the same quality of plan at three in the morning as at three in the afternoon; a human does not. If you have watched a night-shift planner juggle six arrivals, a channel closure and a broken grab, you know that degradation is real and you know nobody measures it. The other is the eight to ten years. That is the number that should make terminal managers sit up. Berth planning capability has been a people pipeline with a decade-long lead time, and it walks out of the door when the person retires. Compressing the rule-heavy bulk of that job into something you can deploy is a supply-chain answer to a labour problem, and the 26 scenarios with 19 agents are the tell: this is being spread thin and wide rather than parked on one berth group.

Now the part I find genuinely interesting, which is where the 10% throughput claim stops being a press number and becomes a queue number. In a queue, waiting time does not fall in proportion to added capacity. It falls faster than that when the system is already busy, and much slower when it is not. The rough shape is waiting proportional to rho over one minus rho, where rho is utilisation. Take a berth group running at rho = 0.85. Its queue factor is 5.67. Add 10% effective capacity and rho falls to about 0.773, queue factor 3.40 — a 40% cut in the waiting term from a 10% capacity gain. Now take a berth group running at rho = 0.50. Its factor is 1.00; add the same 10% and you get 0.455 and a factor of 0.83, a 17% cut. Same model, same 10%, wildly different payoff depending on how congested the berth group was to begin with.

That one curve answers the who-feels-this question better than any stakeholder list would. A congested dry bulk berth group at a big import port is the only place where this pays out materially. A terminal with slack capacity, or one whose real constraint is yard space rather than berths, gets very little — and note that Qingdao's own yard number, re-handling down 5%, is smaller than its berth number, which is consistent with the yard still being the tighter constraint. For the shipowner the benefit lands as vessel-hours at anchor. For the cargo side — the steel mill, the power plant, the cement works taking iron ore, coal and clinker through Qingdao — it lands as a change in the arrival time distribution, and the size of that change is the entire question. For a mid-sized forwarder booking space on somebody else's vessel this is a rounding error, and it stays one until the data starts flowing outward.

Timing is where I would push back on the excitement. The port side gets the benefit the day the agent goes live, and Qingdao has already reported its figures. The cargo side gets nothing until the improved plan shows up inside the ETA that the agent sends you, and that is a data pipeline question, not a model question. My read: inside a single port, six to twelve months before a receiver sees a measurable change in arrival predictability; eighteen to thirty-six months before this is normal across the major Chinese import ports, because each port has to rebuild its own rule set and clean its own data. On the money side, most dry bulk moves under charterparties or annual supply contracts with laytime and demurrage terms, so the first place this reaches your P&L is the next contract year, not this quarter.

Here is the thing I have not seen in any of the coverage, and it is the reason I sat down to write this. The bottleneck moved. It used to be that the plan took two to three hours to produce, so planning was slow and the binding constraint was compute plus human attention. Once the plan takes 90 seconds, the binding constraint moves upstream to the inputs. On a dry bulk berth plan the inputs are two numbers: when the vessel actually arrives, and how fast it will actually discharge.

Let me make that concrete, because bad data is the kind of complaint people nod at and then ignore. Assume — my assumption, not a published figure — that the ETA on a dry bulk vessel 72 hours out carries a mean absolute error of nine hours. That is a normal order of magnitude for tramp bulk; liner services do better, bulk does not. Now put five vessels in a queue. The fifth vessel's slot in your plan depends on the arrival errors of everything ahead of it, and errors of that kind compound with the square root of the number of steps: nine hours times the square root of five is roughly 20 hours. So the ninety-second optimal plan for the fifth berth is optimising against an arrival time that may be off by most of a day. It is a beautiful wrong answer, produced very quickly.

Speed without fresh inputs also creates a failure mode that slow planning never had. If you can re-solve in 90 seconds, you will, and every re-solve driven by a stale or noisy input produces a different plan. The yard has just positioned a grab for vessel A, the plan flips, and now it is vessel B. I have watched this pattern in yard management rollouts and in transport management rollouts, and there is a name for it in the trade: plan churn. Operators stop trusting the plan and start working around it, and real adoption quietly decays while the reported adoption number stays high. Which brings me to the metric.

That taken-more-than-90%-of-the-time figure is the most quoted number in this story and the least informative one. Adoption is not optimality. It means the plan was usable, which is a much weaker claim than it sounds. What I would want instead is the zero-modification adoption rate: how often the plan goes out with nobody touching a field. And I would want the edit log — which fields get changed by hand, and by how much. If planners are accepting the sequence and rewriting the discharge-rate estimate on every single vessel, the model is doing the easy half and the input data is still the bottleneck. Set the target there: move the headline metric from adoption above 90% to zero-modification adoption at 70% within two quarters, and start instrumenting the edit log this month.

Now the money, with the assumptions written out where I made them up. Assume a dry bulk berth group handling four vessel calls a day, 1,460 a year at an average 80,000 tonnes, so roughly 117 million tonnes a year. Assume average waiting at anchor falls from 20 hours to 15 hours. I want to be straight with you that this is my reading and not their number: Qingdao published throughput up 10%, and I am translating that into a waiting-time reduction using the rho curve above, on the optimistic side. Assume vessel time is worth 1,800 dollars an hour, a panamax-class time-charter equivalent plus fuel at anchor, again my assumption. The saving is 1,460 calls times five hours times 1,800, about 13.1 million dollars a year. Spread across 117 million tonnes, that is roughly 0.11 dollars a tonne.

For the receiver, run a second calculation, and this is the more useful one. Assume a mill drawing 120,000 tonnes a month through Qingdao, so about 4,000 tonnes a day of consumption. Assume the standard deviation of vessel arrival is 3.2 days today and falls to 2.0 days. Safety stock against arrival variability is roughly z times sigma; at a 95% service level z is 1.65. Today that is 5.3 days of cover; afterwards 3.3 days. The gap is 1.98 days times 4,000 tonnes, roughly 7,900 tonnes released. At an assumed 100 dollars a tonne that is about 790,000 dollars of working capital freed, and at an assumed 8% annual carrying cost — financing plus storage — roughly 63,000 dollars a year.

Notice what that calculation rewards. It rewards a smaller sigma, not a smaller mean. Five hours off the average waiting time is 0.21 of a day and it does almost nothing to safety stock, because safety stock is held against variability, not against lateness. So the question I would put to Qingdao, or to any port running one of these systems, is not how much faster the plan is. It is whether the variance of arrival time came down, or only the mean. Those two produce completely different balance sheets on the receiving end, and only one of them lets a mill cut inventory.

Where could I be wrong? Two places, and both are testable. The veteran planner's value may sit in exceptions rather than in computation. Weather, a broken conveyor, a pilot shortage, a charterer leaning on the port to jump the queue, the phone call that resolves it — none of that lives in the 180 rules, and it is precisely the part you lose if you let the eight-to-ten-year training pipeline dry up. If the exceptions are where the money is, a model that handles the rule-heavy majority beautifully can still leave the terminal worse off on the bad days, and the bad days are the expensive ones. The other place is the constraint itself. Qingdao's yard number being smaller than its berth number is a hint. If the real queue forms at the yard or at the rail gate, then berth optimisation is polishing a step that is not the bottleneck, and the 10% throughput figure will not reproduce at a port shaped differently.

The test that settles it is cheap. Track mean waiting time at anchor and the 95th percentile waiting time on the same berth group, month by month. If both fall together, the system is genuinely taking load off the queue. If the mean falls and the P95 does not move, what is happening is that the model is picking the easy vessels: the well-behaved arrivals with clean discharge estimates get planned beautifully while the messy tail stays a mess. That outcome is extremely common in optimisation projects and it is invisible unless somebody insists on the percentile.

So what would I do with this, sitting on the cargo side rather than the port side. Start with the baseline, because everything else depends on it. Pull twelve months of your own records and build one table: forecast ETA against actual berthing time, bucketed at 72, 48, 24 and 12 hours out, with mean absolute error and standard deviation in each bucket. You cannot tell whether a 90-second plan helped you until you know how bad your own arrival data was to begin with. I have watched companies layer optimisation on top of a 20-hour ETA error and then wonder why nothing moved. Quantify first, then optimise.

Then get the data calibre written down, and this is the least glamorous and most valuable step of the lot. Who enters the ETA, who updates it, how often, who owns the discharge-rate estimate, and what happens when the two disagree. On a berth plan there are at least four parties with a finger in it — the agent, the terminal planner, the pilot station and the receiver — and in my experience nobody has ever written down which of them is the single source of truth. Set the refresh cadence to match the new solving speed as well: if the model re-solves in 90 seconds, a once-daily update is nonsense. Targets I would hold it to: ETA mean absolute error of 3 hours or better at 48 hours out and 1 hour or better at 12 hours out; discharge-rate estimate within 10% of actual; dynamic fields refreshed every 15 minutes.

And put a guardrail on the re-solve. A plan that can change every 90 seconds will churn unless you bound it. I would require human confirmation whenever a re-solve shifts a vessel by more than one tidal window, or swaps more than two berth assignments against the version the yard is already working to. That single rule costs you almost nothing in plan quality and it protects the operational trust that projects like this die without. Change the metric set while you are at it: waiting time mean plus P95, discharge rate variance, zero-modification adoption, and plan churn events per week. Four numbers, one page, reviewed monthly.

On the commercial side, stop asking your agent for a date and start asking for a distribution. She arrives on the 14th is a number with no error bar attached, and you are setting safety stock against it. Ask for the arrival window and, if they can produce it, the standard deviation they are actually seeing on that service. Then take that sigma into your own inventory formula and work out what your safety stock should really be; most mills I have worked with are carrying cover against a sigma nobody has measured in years. When the contract comes up for renewal, put the P95 waiting time into the laytime conversation rather than the average. Demurrage is driven by the tail, and the tail is the only thing a system like this might genuinely compress.

Leo would tell you this is experience yielding to the machine, and I half agree: the machine wins on consistency and loses on exceptions. What I will not agree with is the headline. Ninety seconds versus two hours is the wrong comparison, because the two hours was never where the money was. The money was always in the queue, and the queue is now being fed by whoever owns the data. Quantify first, then optimise.

↑ Back to the key points

— By Ravi Chandran

port-ailarge-language-modelberth-planningqingdao-port