we want the objection you would raise in a meeting rather than the one you would leave in a comment
we want the objection you would raise in the meeting, not the one you would leave in a comment.
we want the objection you would raise in a meeting rather than the one you would leave in a comment
we want the objection you would raise in the meeting, not the one you would leave in a comment.
It needs the safety engineer’s rigor about what a requirement actually means, the actuary’s care about how much weight a number can carry, and the operator’s pragmatism about what a working floor will really keep.
It needs the safety engineer’s rigor about what a requirement actually means, the actuary’s care about how much weight a number can carry, and the operator’s pragmatism about what a working floor will actually maintain.
A measured section a machine can check is worth more than a paragraph an auditor interprets.
A measurable requirement a machine can check is worth more than a paragraph an auditor has to interpret.
The application-side work item is the gap, and a national member body is who can open it.
The application-side work item is the gap, and a national member body can open it.
Surplus lines is the door this class can come in through, and the standard market is where the record eventually carries it.[46] Whether it lands as endorsements on the lines that already run, or a product of its own, is the record’s call, not ours.
Surplus lines is the door this class can come in through, and the standard market is where it can eventually graduate. [46] Whether that happens through endorsements on existing lines or a product of its own is for the market to decide, not us.
A blanket AI exclusion is one of two available answers.
A blanket AI exclusion is one of two available answers to this risk.
Ops already needs the telemetry; kept to the requirements, it is the book the production deployment arrives with.
Ops already needs the telemetry; turned into a record that meets these requirements, it becomes the book the production deployment arrives with.
Stamp every hour of the record with who was in control: autonomous, supervised, teleoperated, or in handover.
Stamp every hour of the record with the operating mode: autonomous, supervised, teleoperated, or in handover.
The record is the only one of those you control.
The record is the one part you control.
The three after them are requests.
The three sections after them are requests.
If pilots convert to production at scale through 2028 while the evidence stays ad hoc and the exclusions stay in place, then legible risk was never the bottleneck, and we overstated the asymmetry this document rests on. If carriers respond to the record by widening the exclusions rather than pricing on it, and no dedicated product outlives the current specialty programs, then insurance will not play its position in the loop. The record would still exist, but only where a law demands it. That is China’s answer, not a market’s. If claims frequency per robot-hour stays flat as uncaged deployments scale, then robot risk was more legible than we claimed, and the cage-era actuarial tables would have carried it. One caution on the test. Raw incident counts misfire at base rates this small, and the workplace injury rates on the books describe the programmed robot in its cage, not this class. Each test above is stated per robot-hour. Someone has to be keeping the hours.
If pilots convert to production at scale through 2028 while the evidence stays ad hoc and the exclusions stay in place, then legible risk was never the bottleneck. We overstated the asymmetry this document rests on.
If carriers respond to the record by widening exclusions rather than pricing from it, and no dedicated product outlives the current specialty programs, then insurance will not play its role in the loop. The record would still exist, but only where a law demands it. That is China’s answer, not a market’s.
If claims frequency per robot-hour stays flat as uncaged deployments scale, then robot risk was more legible than we claimed, and the cage-era actuarial tables were already good enough to price it.
One caution on the test: raw incident counts misfire when the base rate is this small, and the workplace injury rates on the books describe the programmed robot in its cage, not this class. Each test above is stated per robot-hour. Someone has to be keeping the hours.
Today that evidence is generated ad hoc and thrown away. Evidence from the pilot starts the Evidence Loop.
Today, much of that evidence is generated ad hoc and thrown away.
Evidence from the pilot starts the Evidence Loop.
That is the number the three renewals already run can read: the site’s workers’ comp, the site’s liability cover, and the maker’s product line.
That is the number the three renewals already run can read: the site’s workers’ comp, the site’s liability cover, and the maker’s product line.[49]
The mode split is only the first cut; the hour will divide again — task, payload, how close the people are — the way rating factors divided the mile.
The mode split is only the first cut. Robot-hours will eventually need to be broken down by task, payload, proximity to people, and other factors, just as auto risk eventually broke the mile into rating factors.
The
The denominator also has to respect the fact that different operating modes carry different risks.
It will also be rare for years: forty-one robot-related deaths in twenty-six, and no one could state a rate, because there was no denominator.
The incidents will also be rare for years. There were forty-one robot-related deaths in twenty-six years, but no one could state a meaningful rate because there was no denominator.
The robot that decides has none.
The robot that decides has no equivalent denominator yet.
The
Once the record is consistent, it can produce a number the market can actually use.
Teleoperation is still essential to how these robots run, and will be for a long time. The ops record has to say who was in control: autonomous | supervised | teleoperated | handover-transition. The handover is the sharpest: who held authority, what triggered the transfer, how long to acknowledge. Today, which policy responds has to be fought over claim by claim, and everyone pays for the uncertainty. That field protects the operator, the deployer, and the maker from each other.
Teleoperation is still essential to how these robots run, and will be for a long time. A robot may be autonomous one minute and under human control the next. The ops record needs to capture the operating mode at every point: autonomous, supervised, teleoperated, or transitioning between them.
The most important moment is the handover. The record should show who had control, why the transfer happened, and how long it took.
Today, when something goes wrong, determining which policy responds can become a claim-by-claim dispute, and everyone pays for the uncertainty. A clear record of control makes responsibility easier to establish for the operator, the deployer, and the maker.
The telemetry is not new work. The ops team reads it every day: uptime, interventions, what an update broke. Keeping it to those requirements is new work, and the deployer pays for it. That is the cost of turning telemetry into evidence. The deployer also owns what it pays for: this is the one record the site can take to its own renewal.
The telemetry is not new work. The ops team reads it every day: uptime, interventions, what an update broke. The new work is turning that telemetry into a record that meets those requirements, and the deployer also owns what it pays for. This is the one record the site can take to its own renewal. There is already a familiar model for this: the flight recorder.
The
The data itself is not new. What is new is turning that data into evidence.
A robot works among badge readers, patient areas and product designs, and what the record leaves out is named up front rather than dropped later
A robot works among badge readers, patient areas and product designs. What the record leaves out should be named up front, not discovered later.
Complete is scoped, not total.
“Complete” is scoped, not total.
The
For that record to be useful as evidence, it has to be trustworthy.
A model’s flaw ships to every unit at once, and neither side can assemble the other’s view alone
A model’s flaw can reach every unit at once, while the maker and deployer each see only part of the picture
It adds up two ways: one deployment, for that site’s insurance, and one model version across every site it runs on, for the maker’s.
It adds value in two ways: it gives the site a history for its insurance, and it gives the maker a history of each model version across every site where it runs.
After
But the stamp only tells us where the deployment started. The record tells us what happened after that.
It is rechecked when the model updates, or when the site’s operation does.
It is rechecked when the model updates, or when the site’s operation changes.
Existing standards roll into it, the battery’s cert and the site’s OSHA rules, what already makes the parts and the workplace safe.
Existing standards roll into it: the battery’s cert, the site’s OSHA rules, and whatever already makes the parts and workplace safe.
That
Before the record starts, the deployment still needs its day-one proof.
Only the deployer can sign for the deployment as it runs.[26] One record format covers every form factor because risk is realized at the deployment. The safety case stays specific to the application.
The deployer is the one who can account for what happens on the floor.[26] One record format can cover different robot form factors because risk is realized at the deployment. The safety case stays specific to the application.
The instrument attaches to the deployment, not the robot.
The record has to follow the deployment, not just the machine.
The pieces exist. Nothing binds them as a system.
The pieces are already here. What is missing is the system that connects them.
The record is the one object all of them can read, each for a different decision.
The record is the one thing all of them can use, each to make a different decision.
None of this waits on new authority.
None of this requires waiting for a new authority.
VDA
Even the systems connecting robots to their environments are not yet built to capture that history.
On the maker side, NVIDIA’s accredited lab is preparing integrations for certification by TÜV, UL and exida, and FORT is working on shipping the same layer into fleets.[31] But these certifications cover the design. They say nothing about what the machine did last quarter.
On the maker side, NVIDIA’s accredited lab is preparing integrations for certification by TÜV, UL and exida, and FORT is working on shipping the same layer into fleets.[31] But these certifications cover the design. They say nothing about what the machine did last quarter.
China
Regulation and standards are moving too, but they are solving different pieces of the problem.
Robot cover is at the start of that path: exclusions have landed, and a dedicated product is starting to be written. Relm wraps existing cover; brokers place programs the ordinary market will not take.
Robotics is beginning to follow the same path.
Exclusions have landed, and a dedicated product is starting to be written. Relm wraps existing cover; brokers place programs the ordinary market will not take.
Cyber already ran the Evidence Loop. SOC 2 Type 1 is a snapshot of the controls; Type 2 is whether they held for a period. Security teams made the report a condition of the deal. Cyber insurers closed the circle with instrumentation: they priced the risk, and fed what they saw back so the next attack did not land.[28] AI agents are running the same loop now.[38] The same loop is already running for people at work: priced from the camera footage, and the record prevents the next injury.[39]
The Evidence Loop is not just a theory. Parts of it are already working in other industries, and some are beginning to appear in robotics.
Cybersecurity already runs a version of the Evidence Loop.
SOC 2 Type 1 is a snapshot of the controls; Type 2 is whether they held for a period. Security teams made the report a condition of the deal. Cyber insurers closed the circle with instrumentation: they priced the risk, and fed what they saw back so the next attack did not land.[28] AI agents are running the same loop now.[38]
The lesson is simple: evidence becomes useful when it can change the next decision.
The loop is already turning for cars that decide. Waymo and Swiss Re stated the risk as claims per million miles, against two hundred billion miles of human driving: eighty-eight percent fewer property claims, ninety-two percent fewer injury claims.[25] At twenty-five million miles, it had two injury claims. That operator was priced because ordinary cars already ran this loop, and those human miles were already on the books. The robot that decides has no such book. It has not started.
Cars that decide are already showing what this can look like. Waymo and Swiss Re stated the risk as claims per million miles, against two hundred billion miles of human driving: eighty-eight percent fewer property claims, ninety-two percent fewer injury claims.[25]
At twenty-five million miles, it had two injury claims. The important part is that the risk had a denominator: miles driven. That history gives an insurer something it can compare and price.
The robot that decides has no equivalent book of history yet. It needs its own exposure measure and its own operational record. That is what has to start now.
ASME is that stamp.[7] IIHS is the crash-test rating in the car ad.[27] In both cases the party that paid when the machine failed built the evidence: Hartford’s inspectors, the crash tests insurers still fund. The ops record has to work the same way: an auditable record, not a log the maker wrote.
ASME is that stamp.[7] IIHS is the crash-test rating in the car ad.[27]
In both cases, the people carrying the financial risk had a reason to build the evidence. Hartford's inspectors gathered it for boilers. Insurers still fund crash testing for cars. The ops record has to work the same way: an auditable record, not a log the maker wrote.
Steam ran the Evidence Loop. Cars run it. The names change. The three rails do not.
Steam boilers showed that evidence could make a dangerous machine insurable and scalable. Cars showed that the same system could evolve as the technology changed.
The mechanism is the same: evidence, insurance, standards.
Safeguarding still stops the machine. A protective stop that fires in a tenth of a second is what keeps the person alive, and no record substitutes for it. What the Evidence Loop replaces is the cage’s other job: being the safety case, the standard, and the policy in one, for a robot that decides. The standard can start thin and thicken as the loop turns. The pieces have run for steam, for cars, for flight. They do not yet run for a robot among people
The loop helps manage risk over time. It does not replace the systems that prevent an accident in the first place.
A protective stop that fires in a tenth of a second is what keeps the person alive, and no record substitutes for it. What the Evidence Loop replaces is the cage’s other job: being the safety case, the standard, and the policy in one, for a robot that decides.
The standard can start with what we can measure and become more precise as deployments generate better evidence.
The pieces have run for steam, for cars, for flight. They do not yet run together for a robot among people.
The boiler inspectors ran it. They did not stamp the machine once. The policy put the inspector on the floor. Inspection was the evidence. The evidence priced the policy. What the inspectors learned became the bar, and more boilers went in. Steam scaled because the Evidence Loop kept turning.
The Evidence Loop isn't a new idea.
The boiler industry was already running it more than a century ago. Inspectors collected evidence from machines in service. Insurers used that evidence to price policies. What inspectors learned helped shape the rules, and better rules made it possible to put more boilers to work.
Steam scaled because the Evidence Loop kept turning.
Evidence separates an observed risk from an assumed one, and that separation is what insurance prices. Insurance feeds the next standard. The standard improves compliance and insurability, and creates more deployments. More deployments write more evidence. That is the Evidence Loop in full.
Evidence turns a risk from something we assume into something we can measure. That gives insurers something to price.
Insurance creates an incentive to collect better evidence, and that evidence can improve the standards used to assess the next deployment.
Better standards make robots easier to comply with, insure, and deploy. More deployments create more evidence.
That is the Evidence Loop.
The stamp and the record, kept together, are evidence. Evidence is what the gates need, and what scales deployments.
The stamp and the record, kept together, are evidence. The stamp shows what was approved. The record shows what happened after deployment. Together, they give the market something it can measure, compare, and eventually price.
The first task is safety clear enough to stamp, and a risk clear enough to insure. Until then, the law still asks, insurers exclude what they cannot bound, and makers retain what the market will not take. The stamp gets the machine onto the floor. It cannot follow a machine that decides. An AI-driven machine adds behavior that design analysis alone cannot settle. Its safety case is statistical, and a statistical case is not finished until it has been measured in service: exposure and outcomes. That is why the case needs a record.
Getting a robot onto the floor therefore requires two things: a safety case clear enough to approve and a risk clear enough to insure. Until then, the law still asks for safety, insurers exclude what they cannot price, and makers retain what the market will not take.
But approval has a limit. A stamp tells you what was true on day one. It cannot tell you what happened after the robot went to work.
An AI-driven machine can behave differently across tasks, environments, and software versions. Design analysis can assess what the machine was built to do; only a record can show what it actually did in service.
That is why the case needs a record.
Getting it insured is the other gate. The buyer asks for a certificate of insurance. Insurers are writing AI damage out of general liability. Four thousand of those exclusions were filed in the year to July 2026, and regulators turned down fewer than one in a hundred.[13] One model flaw reaches every machine running it, so a single defect can touch every policy at once. Robotics founders say they “simply cannot get insured.”[14]
Compliance gets the robot onto the floor. Insurance is the other gate.
The buyer asks for a certificate of insurance. Insurers are writing AI damage out of general liability. Four thousand of those exclusions were filed in the year to July 2026, and regulators turned down fewer than one in a hundred. [13]
One model flaw reaches every machine running it, so a single defect can touch every policy at once. Robotics founders say they “simply cannot get insured.”[14]
Thanks to modern factory and labor safety law, a robot working among people has to be compliant from day one. The first gate for that is a checklist or a case someone will sign: safety assurance clear enough to allow the machine onto that floor. The 2025 revision of the rules for industrial robots reaches the robot, the cell, and now the people who run it.[26] For an AI-driven machine the call is slower and harder, because what it does not reach is the decisions inside the machine.
A robot working around people has to be compliant from day one. The first gate is a safety case clear enough to approve the machine for that floor.
The 2025 revision of the rules for industrial robots reaches the robot, the cell, and now the people who run it. [26]
But for an AI-driven machine, design-time assessment cannot fully account for how the machine will make decisions once it is in use.
The same pattern has emerged many times: for electricity, the elevator, the car, the airplane. The question is what it looks like for a robot.
This pattern has worked before, for electricity, elevators, cars, and airplanes. Each time, measurement turned an uncertain risk into something engineers, regulators, and insurers could act on.
The question now is what that system looks like for a robot that can decide.
The company is still there, operating inside one of the world's largest reinsurers.[8] Six hundred inspectors and engineers still carry commissions from the code bodies, and the boiler inspector of 1867 now ships sensors into its own policies.
The system didn't disappear. The company that began with those early boiler inspections still operates today, now using sensors alongside its inspectors. [8]
What made the system work was not the inspection alone. It was what the measurement made possible. One boiler could be inspected. Many boilers could be compared. And once the risks could be placed on the same scale, insurers could price them
The incentives locked together.
Once the risk could be measured, everyone had a reason to act on what the measurements showed.
The fix came from engineers and an insurance man, wielding a financial instrument with teeth. In Hartford, Connecticut, a small club of them had spent years studying why boilers burst. In 1867 they brought to American industry a product Britain had proved a decade earlier: an insurance policy you could only buy by letting their engineers inspect your boiler, regularly. The inspection was the product, and the policy was priced on its report.[6]
The problem was not just that boilers were dangerous. It was that nobody had a reliable way to measure that danger.
The fix came from engineers and an insurance man, wielding a financial instrument with teeth.
In Hartford, Connecticut, a small group had spent years studying why boilers burst.
In 1867, they introduced an insurance policy that came with a condition: their engineers had to inspect the boiler regularly. The inspection was the basis for the policy, and the policy was priced using what the inspection found.[6]
Steam boilers were the physical AI of the nineteenth century: machines of enormous economic power that nobody could see inside. They drove factories, riverboats, and locomotives. And they exploded.
Steam boilers were among the most powerful machines of the nineteenth century and among the most dangerous.
They powered factories, riverboats, and locomotives, but nobody could see inside them. And they exploded.
Design-time methods exist, but nothing measures what the machine actually did once it is in service. So, what replaces the cage? The answer begins with a boiler.
Design specifications can tell us what a machine was built to do. They cannot tell us how it behaved once it was in service. That is the gap we need to close. And the answer begins with a boiler.
The cage did three jobs at once. It was the safety case, the standard, and the insurance policy. Nobody had to measure how the robot would decide, because it never decided anything. For a robot that decides constantly, its behavior itself must now be measured.
The cage made safety easier to prove.
The robot had a defined set of movements, a defined space, and a clear boundary between the machine and the people around it. That made its safety something you could test before it went to work.
But a robot that can change how it behaves from one situation to the next needs something more: a way to measure what it actually does in the real world.
Even the programmed robot is still paying for the cage. A picking robot moves down an empty warehouse aisle at a snail's pace, because nobody can prove it is safe to go faster. One vendor puts it plainly: “you make big fields and you make the robot drive slow.”[3] Where risk cannot be proven, it is paid for in speed. Call it the caution tax.
Even the programmed robot is still paying for the cage.
A picking robot moves down an empty warehouse aisle at a snail's pace, because nobody can prove it is safe to go faster.
One vendor puts it plainly: “you make big fields and you make the robot drive slow.”[3]
Where risk cannot be proven, it is paid for in speed. Call it the caution tax.
They will be every shape, wherever people work, side by side with us, for decades.
They will take many forms and work wherever people work, side by side with us, for decades.
For sixty years, someone programmed every move, and the robot repeated it: walk near and it slows, open the gate and it stops, the same response every time. Safety was a mechanism you could test. That was the programmed robot.
For sixty years, robots worked within a set of rules we could predict. Walk near one and it slows. Open the gate and it stops. Give it a task and it repeats it the same way every time. Safety was a mechanism you could test. That was the programmed robot.
It is an exciting time to be a robot. They are becoming promptable, not programmed. They are stepping onto factory floors, into warehouses, restaurants, and homes. For sixty years, the answer to keeping us safe around robots was a cage: bolt the machine down, fence it off, keep people out. It worked because the robot was blind and repetitive. The cage kept the risk contained. Dangerous inside, safe outside. Today's robots can see, plan, and work on their own, alongside people. The prize is the largest market on earth: labor itself. But what replaces the cage? Every dangerous machine in the past, from the steam boiler to the elevator to the car, won its place the same way. Its risk became legible: measurable, comparable, priceable. Robot risk today is not legible. But a robot among people today must be compliant and insured from day one. Legibility takes instrumentation to measure, standards to compare, and insurance priced on evidence, turning as one circuit: evidence prices insurance, insurance feeds the standard, and the standard writes the next round of evidence. We call it the Evidence Loop. As a system, it has not started. What follows is our answer: the 17 things robot makers, deployers, insurers, and policymakers can do by 2030 to start the loop that replaces the cage. It is published open for review: we invite you to shape it and join its reviewers.
For sixty years, we knew exactly how to keep a robot safe. It came down to one simple idea: keep people out.
A robot used to wait for instructions. It worked because it was predictable, repetitive, and blind. Now imagine one deciding what happens next.
That assumption is now breaking. The robot is leaving the cage. AI is giving it the ability to see, plan, and act without a script or program for every move. every move.
These machines are moving into warehouses, restaurants, factories, and homes, working alongside the people the cage was built to keep away.
But what replaces the cage?
The machines that changed the world all faced a version of the same problem: how do you make something dangerous safe enough to use around people?
Every dangerous machine in the past, from the steam boiler to the elevator to the car, won its place the same way. Their risks became something people could inspect, compare, and eventually price.
Their risk became legible.
Robot risk isn't there yet.
A robot working around people has to be compliant and insured from day one. But a safety assessment can only tell you what was true when the robot was assessed.
It cannot tell you what happened after the machine went to work: how often it stopped, how often a person had to intervene, what changed after an update, or whether the same problem appeared across deployments.
That matters because insurers cannot price what they cannot measure. And without a history of real-world performance, every new deployment starts with uncertainty.
The missing piece is evidence: a record of what the robot actually did in service, measured in a way that different deployments can be compared.
Once that record exists, the pieces can start reinforcing one another.
It can help insurers price risk, give standards bodies better data, and give deployers a track record they can take to their next renewal.
We call this the Evidence Loop.
This article lays out 17 things robot makers, deployers, insurers, standards bodies, and policymakers can do by 2030 to start the loop that replaces the cage.
It is published open for review: we invite you to shape it and join its reviewers.