Mathematical Optimization for Cargo Ships
research.google
research.google
On the terminal side, all I can say is that I am knee-deep in container optimizations right now and it's a damn nightmare. Every terminal does things extremely differently, even if the terminal is owned by the same company. Even the terminology is often different within a company. You optimize for one terminal, and you have to build 80% of it from scratch for the next terminal. Any solution is incredibly difficult to scale.
To set the stage, this is a blue-collar industry, and what happens, so I've been told, is quite common in blue-collar industries. A lot of good ol' boy networking, people getting promoted who have no business being promoted, and as a consequence, a lot of corruption, ineptitude, and inefficiencies. The software in this space just flat out suck and anyone at a FAANG company would be appalled at how we write software. It's literally longshoremen trying to run software businesses.
On the other hand, as a terminal, it's very hard not to make money, so there's a lot of players, and unfortunately a lot of, at least shady deals.
I'm involved with a company that sells to terminals that don't have home-grown software systems. I don't have that much info on terminals that have their own software staff, however, there just isn't that much breakthroughs in the space. It's a conservative industry with a lot of money already, so there traditionally have been little incentive with high risks.
In addition to the aversion to risk, there's a lack of software product and engineering strategy to invest in optimization properly. One major vendor, I was told by a colleague did invest substantially, failed, and "it almost bought the entire company down." My company has dabbled before. We've had a couple of PhDs on staff. We even had a neighbor search solution like in the article developed. It worked in simulation and testing, and failed miserably in real life. The customer didn't take it, and we never had the vision to actually salvage anything from the effort.
There are a few vendors in this space, but again, the landscape is pretty sad. I know of one that are flat out frauds. They exist because the CEOs used to be in the industry and have credibility. They're still around because they're able to con non-technical terminal managers and string them along for a while. A few more maybe have 1 or 2 competent data scientists, but that's not enough. I've worked with one company that is good and is easily the most qualified, but they have never had the marketing and leadership to make themselves the billions they should be worth. They also simultaneously price themselves out of range for most terminals.
The industry is changing, though, because private equity has gotten involved in many companies. They have taken over some internal initiatives and basically dragged many of these good ol' boys against their will, or forced them out. They've also forced companies to partner with tech giants to address these issues. IMO, this is the only way we will see supply chain optimization revolutions any time soon.
In short: It is not good enough. When it is good enough, customers adopt it.
I don't know how many meetings I've been at where the PE team blatantly states over and over, let us help you, tell us what you want to do, tell us what you need, and we completely flub it. Sometimes our CEO is too paranoid, sometimes too prideful, sometimes too oblivious, but mostly just doesn't have the knowledge on how to run a software company.
PE is going to run out of patience soon, and it won't be good for us. There's going to be headlines of how PE gutted us, but in reality, we gutted ourselves.
BTW the current state of English water companies is a not a good advert for private equity.
You have to realize that these terminals are typically a lot smaller in terms of manpower than you might think. A terminal handling 1 million TEU/year may operate with 2-3 crews in a 16 hour/day schedule for handling ships and trucks, plus some extra people handling the empties. Let's say it's 60 dock workers a day. A typically office might only employ about 12 people in overhead: one for customs/quality stuff, one for invoicing and finance, one HR, a bunch of order entry / customer service people, and 2 or 3 yard/berth/stowage planners.
Unless you look at the biggest of biggest terminals, you have to essentially regard them as SMEs in a very non-sexy business, and they treat tech and fundamental research into optimization accordingly.
All of the problems you can neatly divide and name are interrelated in practice, and trying to have algorithms "magically" figure it out for users is pretty much guaranteed to fail. This business eats software hotshots who think they can swoop in and "just" run some algos, "because it's just a TSP / Constraint Solver / [insert your approach here] right?". This business can definitely use bright minds, but start humbly, and start by talking to real world users please.
You don't appear to have any contact details on your profile, but if you ever would like to exchange thoughts with another company in this business, shoot me an e-mail. I think we're doing some interesting stuff for the <1 million TEU/year segment of terminals, especially if they're heavy into multi modal.
Some planning algorithm like this might work, for sure. But onpy if it is included in whatever software being used for this kind of planning. Otherwise, it is a nice academic excercise at best.
- German engineers balked at tightening features on pre-production vehicles because it would prevent their use for recreational purposes - Billions were spent in overtime in a healthcare setting, seemed trivial to improve scheduling to reduce overtime, but there were numerous union constraints and simply not enough hiring supply
I wonder how realistic this google solution is (Houthi missile constraints?). In my experience it's often more valuable to develop solutions that can easily adjust to unplanned changes rather than provably optimal ones
As an employee at a terminal, you have a) no reason to trust some software people who may have never seen one of them steel boxes up close, b) if this works, you or your colleague gets laid off, and c) if it doesn't work, you're now stuck cleaning up the mess the computer will inevitably make for you, which is way less fun than planning things yourself.
The 5%-10% invoices being incorrect is probably on the low end, I'd probably say it's 5-10% in total invoice VALUE. Because of the sheer amount of containers, you need automated invoice calculation. But if you miss an edge case ("We are allowed to invoice a 5 buck surcharge if the container is filled with at least 100kg of explosive goods, AND was loaded during the night shift because dangerous goods handling requirements now require us to have a safety officer on call, which we normally don't have during the night shift") you are leaving serious amounts of money uninvoiced, and you will never know. One of the reasons not that much effort goes into optimization algorithms is that there is so much low hanging fruit left on getting your invoices out correctly...
Also d.) It works sometimes so you get forced to use it and then it gets ransomwared and you're completely unable to work because they laid off so many people you can't fall back to manual operations.
You're also right about the complications and the need to be adaptable to various situations that arise. Some of that just comes from people who are afraid of getting their hands dirty and really understanding the domain they're working on. Operations Researchers and Industrial Engineers are not much different from Computer Scientists: sometimes the desire for a beautiful and simple solution prevents you from building something useful. Sometimes in order to build something useful, your idealized fast-converging LP turns into a QP/MIP bastardized rat king that only converges on an optimal solution half the time, but still gets a better-than-before solution the rest of the time. And some of us really hate that it has to be that way. The rest of us laugh their way to the bank.
From an academic standpoint it seemed a bit to little academic in presentation though, I would like to see the problem formulation! :)
I couldn’t look up his work exactly because our names are highly generic but if you’re interested, I can forward you his work.
Unlikely. I worked in this business and the idea of mathematical optimisation is always the first one to be suggested by someone who's new to this industry. Global shipping companies are clueless about tech and only do what they absolutely must to comply with the regulations, the rest is ignored. Give them an XML Schema for an EDI doc and they will stuff an Excel sheet into a CDATA. They know their game, they wore out the infinite patience of IBM who wanted to convince them to use blockchain to track shipments.
We recently met with a big tech company who we were told to partner with. They start with an angle that made highly suspicious they basically wanted to build a fully automated terminal. When we told them just one fairly mild story of what union issues we deal with, they all gasped and said, "what?" and quickly dropped it.
I agree with you that objective optimization is largely academic. There are always reasons why it's impossible or impractical to adhere to the most efficient versions of standardized processes. Sometimes these are "stupid" (people) reasons, and frequently they're reasonable justifications that accommodate for any number of externalities (weather, downtime, supply chain disruptions, lumpiness in demand signals and forecasts based on seasonality, etc).
That said, it's almost always smart to start with the most efficient version of a process and then add exception-handling, rather than to create a standardized process based on known exceptions. If you let exceptions become the rule, you'll always be operating less efficiently than optimal.
It also ridiculizes my little coding problems x)
I’ll even make a (sincere!) claim that I suspect will inspire a lot of angry attempts at rebuttals:
The 20 foot shipping container has had a greater impact on the world than large language models are likely to ever achieve.
Read the book first and then tell me why you think I’m wrong. (I’m not wrong :) )
(Assuming by LLM's we are broadly referring to AI, rather than LLM's as a specific technology)
Gen Ai has a long way to go and multiple languages and cultures to mesh with before it achieves the ubiquity of a container.
That is a hard claim to verify. One might say impossible. It is because it makes a statement about the future without defining a time interval. The “are likely to ever achieve” part.
Before you start arguing about how little you think LLMs are worth and how great container ships are, that is not the problem. The problem is that to “verify” something with an “ever achieve” you either have to prove hard limits or wait forever. Proving hard limits rigorously is very hard. Waiting forever is impossible.
Sure, if we are going to litigate this on precision of the language use, then this would fail. Since it isn't a legal document but a short sentence on Hacker News. Given an infinite timeline, LLMs can conceivably be more useful than containers.
However, this warning is relatively empty, especially since you have precluded any calculations between LLMs and containers
> Before you start arguing about how little you think LLMs are worth and how great container ships are.
So... what now?
We conclude that the original claim cannot be verified and we reformulate it in a more useful form.
Sometimes the correct answer is "we do not know", and we all should get more comfortable with that. Here it is even worse than that. It is a "we do not know and we will never know".
Or you can say "I think that statement is true." That is entirely up to you. It tells us about your mind state and you are the primary source of information on that.
Saying it is "easy to verify" is not the same kind of thing. It is about the statement itself.
The most substantial claims for LLMs today come from 2 general areas - code copilots (github published, google published), and reduction in time to proficiency for new hires. That said, it may unlock additional capabilities.
Additionally, language processing has a massive gap when it comes to languages that most of humanity speaks. (Gabriel Nichols, CDT paper)
It appears that LLM capability growth is likely to plateau, although this is to be confirmed. https://arxiv.org/abs/2404.04125
Furthermore, having seen GenAI deployments in workplaces, they also suffer from those issues that plague all ML and AI projects.
As initially stated, Cargo containers work across cultures, are standardized across humanity, and underpin all our goods transport.
Given an infinite timeline, it is well likely that cargo containers will outperform LLMs, because the market for pure information will always be at the mercy of physical goods required to generate that information.
So even if LLMs magically improve and become all pervasive gods, they will still need to transfer goods using cargo containers.
Given that this is a HN comment, and not a dissertation, within that social context, I submit that the original postulate, was correct and sufficient.
My Dry Cargo instructor in particular told of going in to SE Asia ports and expecting to be there for 3 weeks unloading & loading break-bulk cargo and getting drunk every night. After the switch to containers, they were lucky to be in port for 20 hours and had no time to go ashore.
Google OR improves existing solutions by 10%-20% utilization which is incredible.
-heavier containers at the bottom for stability
-refrigerated containers on the inside to reduce heat loss
-containers to be unloaded first near the top
etc
And all in 3 dimensions!
But an automated system sounds like it has its own problems, how does it make sure it has proper ventilation and comes on at the right time? Probably I don't want the container to start venting diesel fumes when it's deep in a stack of lego surrounded on all sides.
It has to be automated since it’s a refrigeration unit that needs to know when to turn on the cooler. The control circuit starts up the generator if it senses no power connection at that time, usually off a standard marine lead acid battery.
It should be straight forward to include full wagon power with the upcoming DAC4EU coupling (UIC 552 says that normal coaches are to have an 800A through connection that's regularly fed with 1000V 16.7Hz or 1500V 50Hz; this is single-wire earth(/track)-return).
Keep in mind that unlike North American freight trains, Europe mainline service is almost fully electrified. Thus the reefer would be grid-fed.
I'm asking because I'm looking to propose a change to the current plans for the electric coupler (use near-field RF instead of the currently-preferred single-pair Ethernet or it's fallback powerline; I think 802.11 has suitable PHY options, especially among the OFDM codes if the near-field chamber is dispersive and/or has problematic resonances in the channel), and throwing in "grid power for reefers" with actual numbers from the reefer container industry would be easy (a change to the coupler is a change, and the additional cable through the wagon shouldn't be that extensive either if it's an economical aluminum type; also this would allow passenger coaches that are currently using the pre-DAC4EU hook and chain coupling to be used with DAC4EU rolling stock).
IIRC international reefers require three phase power with a 50 amp breaker so even 800 amps won't get you very far if they all start up at once.
Also, because of the short time requirement, and some extreme low temperature requirements on some cargo, it needs to be in a place where it can be disconnected and very quickly craned off the ship.
The engineer making and breaking the connections will generally have to manually log the time of these actions and the time of the unload. It's all a very interesting and somewhat complicated process.
In the same vein, port /starboard / bow / stern balance.
> -containers to be unloaded first near the top
Containers with the same destination in separate places so that multiple cranes can run in parallel without infringing on each other's work-zones.
Truly an interesting project for people who get stuck trying to optimize too many things at once. :p
It's strange to me that the cargo space has not been aggressively optimized. It's a pretty substantial part of civilization and, I believe, there's definitely some money sloshing around there.
To be fair, a completely optimal solution definitely seems out of reach, but I'm not interested in those.
Just imagine the contingency planning due to restrictions in unloading.
There are also a lot more constraints than weight and balance and they're constantly changing; some may rule out big chunks of the state space which is very good, but it's still a Hard Problem.
But it's not like there aren't people already doing it.
The insight is to see time as a spatial dimension - then a set of shipping tasks - each with a length determined by time to complete the task - can be packed into a set of 1d bins representing boat schedules!
This is variously known as the supply chain optimization problem in logistics, the minimal makespan problem in manufacturing, and the multiprocessor scheduling problem in compsci. All of these problems have been classically formulated as bin packing problems.
Toy model: Single ship + Single container Bin Packing
Here is how it works for a single boat and a single package going between multiple ports. In this case, the problem is equivalent to a shortest path problem on a graph with weighted nodes, and more traditionally presented as a shortest path problem on a graph with weighted edges. I'll show the model and the graph translations!
The model consists of:
1. One bin - a 1 dimensional line representing the utilization of the ship over time
2. Line segments p_1 .. p_k representing the transit time for each of k possible shipping operations (taking the container from some port to some other port). For this model, assume p_1 = p_k = 0, representing the first and final transits
3. A port relation P_ik saying "transit k can immediately follow transit i," or equivalently, "transit i's arrival port is transit k's departure port"
A span is a sequence of line segments respecting the port relations, starting at p_0 and ending at p_k. The problem is to find a makespan - an optimal span.
This is the same as the shortest path on a directed graph G with nodes weighted p_1 .. p_k, where we draw an arrow i -> j iff P_ij.
This graph can also be "dualized" into a directed graph G' where nodes are ports, and arrows between ports are weighted by transit times. This is the most familiar form of this problem.
Define an equivalence relation i ~ j saying transit i and j are equivalent iiff P_ik = P_jk for all k (i and j arrive at the same port). Now give G' a node for each equivalence class [i], which we call ports. Finally, for any pair of ports [i] and [j], add an arrow [i] -> [j] with weight p_k if and only if there exists k such that
1. P_ik. (Transit k leaves from port [i])
2. k is in [j]. (Transit k arrives at port [j])
Now we have the traditional shortest route problem on a graph with weighted edges. However the bin packing model naturally scales to the case where we have many ships, larger cargo capacities, many containers, and complex transit constraints.
In this general case, we regain the geometric packing aspect as well! This is because the ships are represented by 4 dimensional bins, which are packed with containers in x, y, z AND t dimensions! The spatial part of this packing now has to obey efficiency constraints like minimizing the unpacking and reshuffling that happens at each port! Wild, huh?
Exactly.
I was exposed to this sort of conception of the packing problem when implementing Ant Colony Optimization (ACO) for a programming challenge.
https://en.wikipedia.org/wiki/Ant_colony_optimization_algori...
It's very cool nonetheless
Instead start a spreadsheet with all the people who worked on this product, and contact them as soon as Google announces it is shutting the API down. Or you notice a few people chaining their Linkedin status.
Until any of the OR APIs are exposed via Google Cloud, they won't have SLAs or any reliability guarantees and should not be used except academically. Just my $.02.
https://developers.google.com/maps/documentation/transportat...
https://developers.google.com/optimization/service/reference...
Cargo ships do do much of their maintenance underway, but other than that, I'd say the difference is a lot less. (Not that this is a big deal, more pedantic). Cargo ships turn around in ports over a few days, unloading, loading. They may also wait for hours or days for a berth.
https://www.flightradar24.com/data/aircraft/n513dz - Delta A350. Basically running 24/7 other than 3 hour turnarounds at airports.
At the end of the day it's a minor comment, I suppose, not even rising to the level of 'quibble'. :)
Most well supported solvers use a mathematical paradigm to define your problems, which "normal" people would look at and immediately get overwhelmed. There are nice off-the-shelf solutions for common scheduling problems, but unless you used it from the start, every business has some unique wrinkles which make fully adopting one of those solutions too hard. Either your wrinkle is not supported, or you cannot figure out how to wedge it into the tools provided. If you're really inclined to try to solve your scheduling problems better, the courses you'll find usually assume significant prior programming and/or mathematical knowledge.
I think it should be possible to lean on no-code paradigms to build a modelling environment for scheduling problems that is accessible for "normal" people who need more than Excel but can't hire an OR specialist.
I also have access to Google's OR API as a trusted tester and have been toying with the solveShiftScheduling endpoint. Out of the box, I don't think I'll be able to easily represent some of our constraints to work with the API. Also, our management frequently want to change the scheduling behavior, so while I can reformulate the problem to make it work with the API now, I never know what's coming down the pipe.
To give a stupid example, I was once responsible for all the enterprise apps at an F500 used by Finance, HR, and Supply Chain/Procurement. This included our internal travel request tool, which had embedded approval hierarchies based on HR hierarchies & levels -- as one would expect. In the 2008 recession, part of the belt tightening was a new policy that the CFO had to personally approve any international travel requests, no matter who submitted them. Theoretically it was a very easy change, but can you imagine how unpleasant it was to build that logic into the system in a non-destructive way ... but a way that also had exception-handling to deal with times the CFO wasn't personally available.
If their title is X, they must get approval. EXCEPT This one really good sales person. Their title is X but they can fly first class if they want and doesn't need preapproval unless it's 10k where everyone else is 2k.
Everyone flies except that one Software Developer who has Doctor provided note about their flying phobia so they can take the train (Amtrak). If train schedule doesn't line up, they get to show up a day early or stay extra day to accommodate train schedule.
Anything dealing with People quickly becomes madness.
1. People can be very creative in solving their problems with your features (and bugs!), even completely unrelated at first thought. They just have to be (a) observable, (b) speaking the language your users understand, which mathematically oriented or generic packages do not.
2. On the other hand, the quite plausible (to me) approach of "obtain an initial solution — adjust for ad-hoc constraints — reoptimize the rest" constantly fails as "too complex" with users reverting to Excel instead.
However, I still stubbornly believe that mathematical optimization cannot do everything, and we should aim for domain-specific decision-support systems that are primarily manual where optimization is only a part of solution building, and UX really matters in that process.
Go players: Go is so difficult. So many options and constrain blah blah blah
AlphaGo: Oh yah?
(ok, maybe managers of low margin business are, but I digress)
As some other comment mentioned, these systems have been controversial, because some of them are used in ways that don't take into account normal human needs. E.g. some schedule back-to-back shifts, change with little notice, and can't take into account sorts of real life things (like child care) that a human manager might be able to.
Ballpark, optimistically, shore cranes can do 30-50 moves per hour, 2 or 4, maybe 6 cranes per vessel, and you have to unpack shell layer by layer.
* Ultra Large Container Vessel (ULCV): 14,501 TEU and higher
* New Panamax: 10,000-14,500 TEU
* Post-Panamax: 5,101-10,000 TEU
* Panamax: 3,001-5,100 TEU
24,000 TEU (Twenty-foot Equivalent Units), say 12k 40-foot containers
4 cranes * 50 containers/crane-hour * 24 hours/day = 1.2k containers / day
https://en.wikipedia.org/wiki/Stowage_plan_for_container_shi...Stowage plans for ships also have weight, balance, power, and value acceptability criteria beyond availability at a port.
These overheads made me curious enough to write up some napkin math, since they mention cut-and-run early departures from ports.
As somebody actively involved in the space, there are lots of ways you can make life easier for yourself. Planning in blocks per hatch cover, grouping containers by destination, size, and weight and treating those as mostly interchangeable are table stakes. You then want to send a plan from the ship to the terminal before start of operations, so the terminal can also optimize and shuffle given they know the position of containers in the yard.
Stowage planning is a lot easier if you decide you're going to ignore the details of each container and only occupy yourself with the groups. The end result is very similar for way less work, and you give a lot more flexibility to the terminal to optimize operations.
I wonder why Google didn't just go with an off the shelf solution and integrate it instead of building their own solution?
[1] https://www.optaplanner.org/ [2] https://www.minizinc.org/
Google developed its solver a while ago [1] and it has been open sourced more recently. Additionally, it is also a supported solver for Minizinc [2] so you can use it with a well known tool/syntax without having to rely on the python/C++ libraries. However, those give you access to some of the solvers specific features that can help you speed up solving time.
[1] https://developers.google.com/optimization/ [2] https://www.minizinc.org/doc-2.4.3/en/solvers.html#or-tools
[1] - https://timefold.ai/ [2] - https://timefold.ai/blog/red-hat-optaplanner-end-of-life-not... [3] - https://docs.timefold.ai/timefold-solver/latest/quickstart/o... [4] - https://github.com/TimefoldAI/timefold-quickstarts/
Google Cloud is also building industry specific solutions. Showing you are a leader of research will be helpful in attracting customers.
Who are they competing with?
If they're really improving the state-of-the-art, to the degree that real money is saved, and the industries are big enough, these APIs will probably get used a lot and moved out of research or licensed.
But the question here is why any of the shipping companies should adopt this given the fact that Google likes to deprecate things at a whim?
Shipping companies mostly do this internally already
Aka give me one OR guy and a gurobi license and we will beat their result.
Bonus at the end of the excercise you will have an or guy that can help you solve other use cases as well.
I benchmarked OR-Tools extensively vs e.g. LKH3 + my own hand-written solvers ... it's not a remotely credible library for these things, last time I checked. So ... don't have a lot of faith in Google's offerings here.
http://vrp.galgos.inf.puc-rio.br/index.php/en/updates - where's Google?
And according to the PR they did not use gurobi (or any sota mip solver), just their in-house lp solver in combination with shortest path algorithms and the fix-and-optimize heuristic.
Likely your particular use case will not fit their model anyway (for example you may want to consider contracts of affreighent along with leasing entire vessels).
So yes, or scientist and gurobi will definitely win.
This is just a PR article, it's hard to say what they did or didn't try before arriving at this solution.
Academics care about proving optimality. Businesses care about their own specific use case and not synthetic benchmarks.
This does not align with my experience of people working in OR.