Restoration estimate
Output · R1Cascade Property Group · IM-84213 · fire & water
Verity is the AI copilot that turns a finished 3D scan into a verified property record: room-by-room measurements, materials, damage findings and a restoration estimate, where every number is drafted by the model and signed by a human before it ships.
The record workspace: a dark evidence twin on the left, a signable legal record on the right. Blue is machine-drafted, black is human-signed.
4:40pm. Marcus, a restoration estimator, opens a delivered scan of a fire-and-water loss for Cascade Property Group. The AI has already drafted the whole estimate: rooms, measurements, damage, a dollar total. It looks finished.
But the number he'll defend to a carrier rests on measurements he didn't take, and he can't tell what the software guessed from what it actually saw. So he does what he always does: opens the twin, re-measures three rooms by hand, and the hour the AI was supposed to save is gone.
The tool had answered every question except the only one that mattered: which of these numbers can I actually stake my name on?
InsideMaps captures beautiful spatial data. But restoration and insurance don't buy models, they buy measurements and scopes that survive a dispute. The value, and the risk, lives in the numbers derived from the scan, not the scan itself.
The first attempt at property intelligence was a flat table of extracted fields: no confidence, no provenance, everything presented as equally certain. It looked like automation and behaved like a liability.
Field-shadowing an estimator surfaced the real dynamic: distrust isn't about how often the AI is wrong, it's about not being able to tell when. Three problems compounded.
Analysts don't distrust AI for being wrong. They distrust it because they can't tell WHEN it's wrong, so they re-check everything and the saved time evaporates.
With no link between a number and the scan region behind it, fixing a value meant re-measuring by hand in the viewer. The AI made work instead of removing it.
When a carrier or owner disputes an estimate, there's nothing to stand behind without a record of what the AI proposed versus what a human finalized. Accuracy without an audit trail isn't enough.
A raw number is only a guess. Verity also carries where it came from, and who signed it.
The same measurement is a guess, a location, and a signed fact at once. The raw export only ever showed the first. Verity's job is to show all three, and make getting from the first to the third nearly free.
Before designing a screen, I sat with estimators scoping real water and fire losses, then watched them re-key every measurement into Xactimate at the desk. The 3D scan had already saved the hour of measuring. They spent it again by hand, because a flat AI table gives you no way to tell a fact from a confident guess.
Restoration estimators, plus 3 carrier adjusters, on how a scoped number gets disputed
Scoping a live loss on-site, then re-keying it into Xactimate back at the desk
Where estimators trust an auto-measurement and where they still re-check by tape
Exported scan tables audited line-by-line against the sketches they replaced
Estimators don't need fewer AI errors, they need to see which value might be one, so calibrated confidence rides on every number instead of the file as a whole.
A flat table makes a confident wrong number cost more than a blank one, so machine-drafted values stay visibly drafted in blue until a human signs them in black.
The value gets disputed by an adjuster months later, so provenance and an audit trail have to be built at draft time, not reconstructed under challenge.
Re-keying the scan into Xactimate erases the scan's time savings, so the verified record has to be the deliverable, not a source to copy from.
Design-target archetypes, drawn from the interviews, ride-alongs and the problem space.
Verity is not an extractor. It is a trust instrument. Its job is to make uncertainty legible and correction nearly free, so a human's attention goes only to the values that actually need judgment.
That reframe set two non-negotiables. Calibrated confidence must be shown, never hidden. And provenance must link every value to the region that produced it. Everything else, down to the color law, serves the moment an estimator trusts the record enough to stop re-measuring.
AI is probabilistic and will be wrong. Its output feeds a regulated, disputable document. One estimator uses this in long focused sessions, working from a single existing capture. Five principles fell out of that box.
A calibrated per-value score is always on screen. Hiding uncertainty to look clean is a lie the user pays for later.
Every derived value links to the region that produced it. An unlinked number can't be corrected fast or defended at all.
The edit path is keyboard-first and inline. Friction here is fatal: a correction flow that fights the analyst's speed gets its defaults slammed.
One action logs old-to-new for the dispute file and feeds the calibration loop. The correction is never thrown away.
Auto-accept is earned per field type against a real reliability curve, not granted by a hardcoded threshold that flatters the demo.
Judgment shows in what you refuse. Each of these was tempting, and each would have quietly broken the thesis.
A full immersive 3D walkthrough as the primary surface. Tempting: it demos beautifully, it's unmistakably spatial, and InsideMaps already has the captures.
Why I killed it: The job is verifying numbers, not flying through space. I demoted the 3D to a focused evidence pane that answers one question: where did this value come from? Spatial data became a signed document, not a video game.
Let the model write the whole estimate and skip the human. Tempting: it's the cleanest leverage story, and it gets an AI project funded.
Why I killed it: It destroys the one thing a restoration estimate rests on: when an auto-accepted number is disputed, no one can stand behind it. I kept the human on the values that carry money and risk, and automated only the ones the model earned.
A clean conversational UI over the scan: type a question, get a scope. Tempting: it feels modern and effortless.
Why I killed it: There's nothing to verify against. Without the spatial record beside the answer you're trusting prose, with no way to correct a value or sign it. I made Q&A a grounded, cited layer on the verified record, not a replacement for it.
Hide the messy uncertainty so the product feels calm and trustworthy. Tempting: uncertainty looks like a flaw in a demo.
Why I killed it: It felt more trustworthy while being less so: a pretty lie. I killed it on ethics and calibration. The honest move is to show the score, then earn trust by making it accurate.
Verity is grounded in a real object model, so the audit trail is a first-class object rather than a log line, and every value knows where it came from.
A Provenance link points every Measurement, DamageFinding and MaterialDetection at a SourceRegion (a 3D bounding box or image crop, plus model, version and confidence). A Verification event records who, when, and old-to-new value. That event is the audit trail and the training signal at once.
Verity is human-in-the-loop by construction: the machine drafts every value in blue, Marcus signs only the ones he will defend in black. These are the two paths he actually walks at his desk, one to clear the record value by value, one to turn damage findings into a figure a carrier cannot wave off.
Where PulseOps spent color on operational status, Verity spends it on one question, has a human signed this, plus a separate scarce budget for how sure the model is. The two systems never collide.
The record workspace carries the whole trust idea, so it took the most iteration. It began as a wall of five paper screens, each pinned to one open question; the greyscale pass locked confidence and provenance before color could flatter a weak decision; the shipped screen gave it the dark-evidence, light-record voice, where machine-blue values settle to signed black.
The first wall wasn't about layout, it was about the open question every screen had to answer: where does a value's evidence live, and how does a human sign it? Five sketches, one margin note each.
Greyscale forced the hard calls before color could flatter them: a confidence ring on every value, focus that flies the evidence pane to its source region, and a machine-versus-human authorship split. Annotated for the ML lead and the engineering handoff.
The shipped workspace gives that structure its voice: a dark evidence twin on the left, a light legal record on the right, machine-blue values settling to signed black as the estimator verifies. The color law does the trust work the wireframe only promised.
The AI drafts the whole record in seconds. The design's job is the rest: show sixty values without overwhelm, let a human verify one at a glance, and make every correction teach the system. Focusing a value flies the evidence pane to its source; A accepts, E edits, and the value flips from machine-blue to human-black.
The domain-deep surface. AI-detected damage glows in the evidence pane as translucent regions and numbered pins, each carrying a defect code, severity, and confidence. The estimator walks the pins with J and K, adjudicates each finding, and a confirmed finding maps straight to estimate line items.
The estimate ledger is the outward-facing deliverable. Every line item links back to a measurement or a damage finding, the total splits into dollars a human confirmed versus dollars still AI-estimated, and export carries the full AI-proposed versus human-final audit trail.
Verity's answer to a capacity layer, but for model trust. A reliability curve asks: when the model says 90%, is it right 90% of the time? A movable auto-accept line, drawn against that curve, lets a human decide how aggressive the automation is allowed to be.
A natural-language panel grounded in the verified record. Before answering, it shows which rooms and measurements it will read. Every answer carries inline citations that fly the camera to the cited region, and it refuses when the evidence isn't there.
Every principle above was forged by a failure below. This is where the design actually came from.
The first record hid confidence to look clean and premium. Every value rendered the same way: calm and certain.
The clean look made a low-confidence guess read as fact. A vaulted-ceiling height, the value the model is least sure of, rendered identically to a floor area it nails, so nothing told the analyst to check it. The interface had laundered a guess into a number.
Confidence has to be visible AND calibrated. Visibility alone adds anxiety; calibration alone is invisible. You need both. This forged the show-confidence principle.
Always-on confidence rings on every value, a machine-ink to human-ink authorship law, and a dedicated calibration layer so the score is honest, not decorative.
Correction opened a modal form with fields for the new value, a reason, and a note. Structured data, in theory.
The modal cost about eleven seconds an edit, and a path that slow doesn't survive real use: analysts batch corrections, then skip them and accept defaults to move on. The data would be slower AND worse.
A correction path that fights the analyst's speed will lose, and take the audit trail down with it. The signal has to ride along with the decision, not gate it.
Inline one-key edit: E opens a number field with the dimension line drawn on the geometry, the change logs old-to-new automatically, and the record climbs toward verified without a single modal.
A single fixed 80% auto-accept line applied to every field type at once. Simple, legible, one number.
It over-accepted vaulted-ceiling heights, the field the model is weakest on, while needlessly holding back floor areas it nails. One threshold can't fit every field.
Auto-accept is a per-field-type decision against real reliability, and how aggressive it gets should be a control the human owns.
Per-field-type calibrated thresholds plus a movable auto-accept line drawn against the reliability curve, trading throughput against review load in the open.
Two passes pressure-test whether the trust holds. First I ran the working prototype past domain experts and sat a working restoration estimator through it think-aloud on a real delivered scan; that pass is done and already moved the design. Then I scoped a moderated study to put Verity in front of estimators who have never seen it, protocol below, ready to run.
Walked the live prototype through design reviews with a restoration domain expert and a claims-side adjuster, then sat a working estimator through the record workspace on a real delivered scan, think-aloud. Every place he reached for his tape instead of the screen became a to-do. Three of them moved the design:
Wiring provenance onto the confidence ring was the moment it clicked: a low-confidence value stopped meaning re-measure and started meaning a ten-second look at the source, which is the entire time-saving the scan had been promising and never delivering.
Will an estimator stop hand-re-measuring the values Verity marks high-confidence, and still catch the ones it marks low, well enough to sign a number he would defend to a carrier?
The damage taxonomy was co-defined with estimators; where confidence comes from was negotiated with the ML lead; the MVP scope was cut with the PM.
Accepted higher compute and storage cost to link every value to a source region, because an unlinked number can't be corrected quickly or defended at all.
Accepted more human review to buy fewer shipped errors. In restoration an error becomes a dispute, so a little slack is far cheaper than a wrong number.
Sequenced the field mobile companion to v2. The highest-stakes work happens at the desk where the estimate is signed and sent.
These are design targets, not shipped results. Each pairs a leading mechanism I can point to in the interface with the lagging outcome it is meant to move, and every figure is tagged TARGET.
Every value carries its source region, so verifying is a glance and one key rather than a manual re-measure. Measured from record open to fully signed.
Needs-review sorts to the top and calibrated-high collapses into quiet rows, so attention lands only on the roughly one-fifth of fields that need judgment.
The model handles the fields it is provably good at, so humans spend attention where it actually moves the number and the risk.
AI-proposed versus human-final is logged per field, so a disputed estimate is defensible in minutes instead of a re-measure and an argument.
Every human edit feeds calibration, so the tool gets less wrong the more it is used. This is the flywheel that makes the product compound.
Validation method: log every AI-proposed versus human-final value and measure correction rate, calibration drift and time-per-record from telemetry, in a phased rollout that controls for confounders.
The honest risk is that a confidence number is only as good as its calibration, so a well-displayed but miscalibrated score is the real product danger. I would validate first whether showing confidence actually calibrates trust or merely adds anxiety.
From enterprise teams to growing startups.