InsideMaps · Operations Command Center

Running a nationwide 3D-capture network in real time

PulseOps is the internal command center for running InsideMaps' nationwide 3D-capture network: a live capture-to-delivery pipeline, a 3D-model QA workflow, fleet utilization, and multi-tenant SLA analytics. Designed to double delivery volume without doubling the humans watching it.

RoleTechnical UX Designer
PlatformInternal web · operator workstations
At a glance
The problemCoordinators rebuilt the state of a nationwide capture network from three tools every shift, with no view of what was about to breach SLA until it already had.
The outcomeDesigned and handed to engineering for implementation. Target: time-to-triage from 15–20 min to under 2, at 2–3× volume without adding coordinators.
My roleSole designer, with a PM, eng lead, and ops lead.
TimelineInsideMaps · 2021
ApproachA manage-by-exception command center: one risk-ranked triage queue, per-tenant health, and calm real-time.
Capture Pipeline
77 active jobs
Updated 3s ago
⌘K
6
Filtersstatus: active ×SLA: at-risk ×type: residential ×77 jobs sorted most-at-risk first
Capture12
IM-8428018h 40m
Meridian Homes
2,100 sqft · 8mD
IM-842766h 12m
Redwood Realty
3,400 sqft · 17mM
IM-842713h 30m
Sentinel Insurance
1,850 sqft · 31mD
+ 9 more
Upload9
IM-842582h 48m
Cascade Property
5,200 sqft · 12mP
IM-842495h 10m
Meridian Homes
2,700 sqft · 5mD
+ 7 more
Processing24WIP limit 20
IM-842132h 14m
Redwood Realty
4,200 sqft · 41mM
IM-839982h 14m
Cascade Property
6,050 sqft · 38mP
IM-841904h 02m
Sentinel Insurance
3,300 sqft · 22mD
IM-841767h 20m
Meridian Homes
1,900 sqft · 9mD
+ 20 more
QA183 breaching
IM-841090h 52m
Redwood Realty
3,800 sqft · 1h 12mP
IM-840771h 36m
Cascade Property
2,600 sqft · 48mP
IM-840513h 10m
Sentinel Insurance
4,900 sqft · 26mM
+ 15 more
Delivered214
IM-84002delivered
Meridian Homes
2,300 sqft · on timeD
IM-83981delivered
Redwood Realty
3,100 sqft · on timeM
+ 212 more
Flow · arrivals vs completions todayProcessing net +9 backlog
Capture
Upload
Processing
QA
Delivered

The live pipeline board: every job flowing capture to delivered, sorted most-at-risk first.

The 90 minutes that started this

A job went late while everyone was watching

2:10pm. Renata, a coordinator, is toggling between three browser tabs. A Cascade Property Group order is ninety minutes from its SLA and she does not know it. The job finished processing at 11am and has sat in an unwatched QA queue ever since.

She finds out the way she always does: the account manager pastes a red-faced client message into Slack. The truck already went home. A recapture now means the deliverable is a day late at twice the cost. Nobody did anything wrong. The tools simply could not show what was stuck.

The failure was not a person. It was that no screen in the company could answer one question: what is about to be late, right now?

Business stakes

What it takes to run a capture network

Operationally, InsideMaps is a logistics plus GPU-compute plus human-QA machine wrapped in SaaS. A client orders a digital twin; a technician with an iPhone LiDAR rig captures it; a compute pipeline reconstructs it; a human confirms the measurements; and it must deliver inside a contractual window. Every handoff leaks time, quality, or money.

24–48hContractual turnaround, tighter on enterprise rush tiers
~35 miCapture radius per technician: the network is a patchwork of local zones
~12%Of project value lost to rework when a capture is poor
~2×Cost of a recapture: the truck rolls again and the SLA is blown
Before PulseOps

Three tools, none of which could see a deadline

The ops team ran the network on a legacy admin panel, a shared spreadsheet, and a wall of Slack channels. The spreadsheet was the real source of truth, which is the tell: the company was paying humans to be a database sync process, and by mid-afternoon the sheet and reality had drifted two to three hours apart.

insidemapsOrder AdminA
San Jose — Orders(market: San Jose only)
Order ID
Customer
Address
Status
Scanner
Created
Due date
Actions
IM-84213
Redwood Realty
1428 Camino Real
Processing
mrodriguez
06/28/2020
6/29/2020
Edit
IM-84109
Redwood Realty
88 Winchester Blvd
In QA
pshah
6/28/20
6/29
Edit
IM-83998
Cascade Property
2100 Broadway
processing
pshah
6/27/2020
6/28/2020
Edit
IM-84077
Meridian Homes
512 Alameda
Processing
dana.p
6/28/20
6/30/2020
Edit
IM-84051
Sentinel Insurance
9 Mission St
scheduled
mrodriguez
5/7/21
Edit
IM-84044
Cascade Property
41 K Street
PROC.
sam.t
06/28/2020
6/29/2020
Edit
IM-84012
Meridian Homes
512 Alameda
Delivered
dana.p
6/27/20
6/29/2020
Edit
IM-83981
Redwood Realty
77 Park Ave
delivered
mrodriguez
6/26/2020
6/28/2020
Edit
Rows per page: 10 ▾1–10 of 2,481‹ ›
The Admin: a stock auto-generated CRUD table. Free-text colored status, mixed date formats, a bare due date with no clock, and a detail panel built to edit a record, not resolve a problem.
SLA_TRACKER_v7_FINAL_use-this-one
File Edit View Insert Format Data Tools Extensions Help
DM
G3ƒx=TEXT(F3-NOW(),"h") → #REF!
MASTER — do not edit without asking Dana. do NOT sort (breaks the HRS LEFT formulas). West tab is ~2-3h behind by afternoon.
·
A
B
C
D
E
F
G
H
I
J
K
1
Job ID
Client
Market
Tech
Ordered
Due
HRS LEFT
Stage (typed)
Last checked
PROBLEM?
Notes
2
IM-84213
Redwood
San Jose
Marcus
06/28/2020
6/29 2pm
#REF!
Processing
11:04a
??
system still says Processing but AM says it's late
3
IM-84109
Redwood
San Jose
Priya
6/28/20
6/29
0.9
in QA
10:40a
yes
blurry, might need recapture
4
IM-83998
Cascade
Oakland
Priya
6/27/2020
6/28 EOD
-1.2
proc.
9:15a
YES
LATE. client emailed. escalated to Sam
5
IM-84077
Meridian
San Jose
Dana
6/28/20
6/30
18.4
PROCESSING
8:50a
no
6
IM-84051
Sentinel
Oakland
Marcus
5/7/21
?
#VALUE!
scheduled
?
no due date in contract PDF, asked account mgr
7
IM-84044
Cascade
Sacramento
Sam
06/28/2020
6/29 10a
2.1
capture?
9:55a
?
called tech, no answer x2
8
IM-84012
Meridian
San Jose
Dana
6/27/20
6/29
0
DONE?
yest.
no
delivered I think, double check
9
IM-83981
Redwood
San Jose
Marcus
6/26/2020
6/28
done
delivered
6/28
no
10
IM-83975
Sentinel
Oakland
Leo
6/28/20
6/30 noon
26.5
upload
10:20a
no
11
12
13
West (MASTER)MiamiNashvilleCharlotteDO NOT DELETEoldSheet14
The coordinator's private spreadsheet: hand-typed status, hand-maintained hours-left riddled with #REF! errors, conditional formatting with no legend, a banner begging no one to sort it. A human rebuilt, by hand, the tool the software should have been.
Every later decision pays back one of these six failures
  1. 1Status was arbitrary color words with no time dimension.
  2. 2One word, “Processing,” hid capture, upload, compute and QA.
  3. 3There was no SLA anywhere in the interface.
  4. 4Every view was a single-market silo.
  5. 5It was built to edit a record, not resolve a problem.
  6. 6The answer to every hard question was “export to a spreadsheet.”
The real problem

The problem was invisible wait time, not a cluttered dashboard

I shadowed coordinators across shifts and counted the context-switches. The private spreadsheet existed only because the panel could not show what was stuck, and every shift began with fifteen to twenty minutes rebuilding where the network stood. Three failures compounded.

  1. 01

    Work stalled invisibly between stages

    A job finished upload and waited on a free GPU slot; finished processing and waited on a free QA reviewer. Most elapsed time was wait time, not touch time, and no tool measured where jobs sat.

  2. 02

    Exceptions surfaced after the SLA was already blown

    Failed uploads, compute crashes, no-shows, blurry captures: all discovered reactively, often by a client complaint. Easy jobs got cleared while at-risk jobs quietly aged past their deadline.

  3. 03

    Coordination scaled linearly with volume

    Every job needed manual watching, status chasing, and hand assignment. That is the exact constraint that stops a B2B ops org from growing revenue without growing headcount in lockstep.

In the legacy Admin
OrderCustomerStatusDue
IM-84213RedwoodProcessing6/29
No time dimension.
Open the record and do the math yourself.
In PulseOps
Ordered41h ago
Captured38h ago
Uploaded36h ago
Processingstuck 34h · COMPUTE_FAIL 6m ago
QAblocked
Deliverydue in 2h 14m

The data always existed. The legacy tool just could not show it in time to act.

Research & Discovery

The network kept no clock, so I timed it by hand

No tool at InsideMaps recorded when a job entered or left a stage, so "slow turnaround" was a feeling nobody could locate. I spent discovery on the operator workstations themselves: shadowing full shifts, interviewing every role that touches a job, and reconstructing real stage timings from the raw order log. One pattern held across all of it, and it reframed the project from "clean up the dashboard" to "make the network watch itself."

shadowing
6 shifts

Sat beside coordinators through full 8-hour shifts in three markets, tallying every tab-toggle and Slack check that rebuilt the network's state.

interviews
11 sessions

One-on-one across all five roles: coordinators, QA reviewers, a reconstruction engineer, the ops lead, and two account managers who own the QBR.

analytics
18 months

Reconstructed real per-stage dwell times from the raw order log, since no tool had ever recorded when a job actually moved between stages.

workshop
2 sessions

Mapped the true pipeline states and defect codes with ops and eng, splitting one word, "Processing," into capture, upload, compute, and QA.

What discovery surfaced
01

Most of the clock is wait time

Jobs spent hours queued between stages, not being worked, and "Processing" hid four different waits, so the target was dwell time, not touch time.

02

Every shift began from zero

With no time dimension in status, coordinators spent 15-20 minutes each shift rebuilding what was urgent before they could act on anything.

03

Alarms arrived as complaints

Exceptions were discovered reactively through a client's Slack message, always after the SLA had already been lost and the truck had gone home.

04

The average hid the failing account

A healthy global SLA number masked individual enterprise tenants quietly breaching, and per-market silos hid the breach from the account managers who owned it.

Affinity map · from raw notes to five themes
PulseOps · Discovery synthesisEdited 2h ago
GARK+3
Share
The invisible wait
“Processing could mean four totally different things”
Pings the compute eng to ask if stuck
No screen shows where a job sits
Easy jobs clear while at-risk ones age
Re-checks the same job every twenty minutes
Put time-in-stage on every card
Insight

Wait time was untracked and unlabeled, so every stage needs a visible time-in-stage clock and aging heat.

Rebuilding state every shift
“First I figure out what's on fire”
Rebuilds the picture from three tabs at login
Status is a color word, not a clock
Twenty minutes gone before the first action
Keeps a private sheet the software should be
Land on what's about to breach, ranked
Insight

Coordinators reconstructed urgency by hand each shift, so the tool must land them on what is breaching.

Alarms arrive as complaints
“I find out when the client's already angry”
Hears of a no-show from Slack, not the tool
Failures surface after the SLA is gone
Truck's already home when we catch it
Catch the exception before it breaches
One ranked queue, deduped by root cause
Insight

Exceptions surfaced only after breach, so detection has to move upstream and rank itself by root cause.

QA under the clock
“Reject reasons live in my head, not data”
Types free-text notes nobody can aggregate
Fix-or-recapture is a money call made blind
A slow rubric just gets defaulted through
Keyboards through the queue to keep pace
One defect code feeds the whole loop
Insight

Review had to stay fast while producing analyzable reasons, so one defect code rides with the decision.

Numbers that can't explain themselves
“We're at 88, so why is Cascade furious?”
Exports to a spreadsheet to answer anything
A global average hides a breaching tenant
Every market is its own blind silo
Track SLA health per tenant, not overall
Attribute the slowdown to a single stage
Insight

Reporting could show a slowdown but never attribute it, so every metric must drill to a stage and a tenant.

Gazi
you
AFFINITY MAP · 5 THEMES
Legend
Voice of user
Observed behavior
Pain point
Opportunity
Affinity map · shift-floor notes clustered into five themes
100%+
Design-target personas
RV
Renata Vaughn
Ops coordinator · protects the SLA clock
What they need

A ranked landing surface that says what to touch first, right now.

Goals
  • See what is about to breach before it does
  • Clear the shift's fires without chasing three tools
  • Spend attention only on jobs that need judgment
Frustrations
  • Every shift starts by rebuilding state by hand
  • Finds out a job is late from an angry Slack
TM
Theo Marsh
QA reviewer · accept, fix, or recapture
What they need

A keyboard-first decision where one defect code carries the signal.

Goals
  • Judge a 3D model fast without cutting corners
  • Make the fix-vs-recapture call with cost in view
  • Leave a reason the rest of the system can use
Frustrations
  • A heavy rubric just gets defaulted through
  • Reject notes are free text nobody can aggregate
DO
Dana Okafor
Ops lead · owns SLA and throughput
What they need

Drillable metrics that pin a slowdown to a stage and a tenant.

Goals
  • Double delivery volume without doubling coordinators
  • Point at the real bottleneck, not a gut feeling
  • Defend enterprise SLAs per account, not on average
Frustrations
  • A healthy global average hides a failing tenant
  • Every answer ends in export to a spreadsheet

Design-target archetypes, drawn from the interviews, ride-alongs and the problem space.

The reframe

The problem was never a cluttered dashboard. It was invisible wait time and reactive triage. So the goal was not to show more; it was to make the network watch itself, and interrupt a human only for the ~10% of jobs that need judgment.

Manage-by-exception is not a UI pattern here. It is the business model: the only thing that breaks the linear-headcount curse. And because a capture network is a queueing system, utilization has a ceiling on purpose: the tool has to refuse to reward over-driving.

Constraints & principles

The box I designed inside

Real-time data at scale on a small eng budget. Eight-hour operator use where density is the point. Multi-tenant data isolation. A legacy data model I could not rewrite, so the UI state machine had to mirror real backend states. And a physical, geo-bounded capture reality. Five principles came out of that box, each a tradeoff, two of them forged by failures later on this page.

01

Density with hierarchy, not minimalism

Operators live in this tool for eight-hour shifts. Information density is a feature. The craft is managing it with a strict altitude system, not hiding it behind whitespace.

02

Exceptions surface themselves

Healthy work collapses into count chips. Only the anomalous earns its own row. The system watches every job so a human only touches the ~10% that need judgment.

03

One source of truth, five lenses

Pipeline, QA, Fleet, Analytics and Client health are views over one shared object model, not five apps. Role sets the default lens and write-permissions, never siloes the data.

04

Every number is drillable

No metric is a dead end. Every aggregate links to its constituents: tile to list to record. A KPI dip is always two clicks from its root cause.

05

Utilization has a ceiling on purpose

A capture network is a queueing system. Past ~85% technician or QA utilization, variance explodes and recapture rises. Slack is designed in, not squeezed out.

Rejected directions

Four tempting versions I did not build

The most important judgment in a project is often what you refuse to ship. The dark command-center goes first, because killing it is why the tool you are looking at is white.

Rejected · dark command center
Command Center
Capture12
IM-842002:14
IM-841994:02
IM-841984:02
Processing24
IM-842002:14
IM-841994:02
IM-841984:02
QA18
IM-842002:14
IM-841994:02
IM-841984:02
Delivered214
IM-842002:14
IM-841994:02
IM-841984:02
Rejected direction 01
01

The dark “mission control” command center

A dark slate canvas with glowing blue and saturated status, styled like a NOC wallboard. Tempting: “command center” pulls you there, it looks impressive in a thumbnail, and there is a real wall display in the ops room.

Why I killed it. Operators do not glance at a wall; they live in this tool for eight hours. If everything glows, nothing is an alarm, which directly undercuts the manage-by-exception thesis. I moved to a white-based system and spent color only on status.

02

Separate role-specific apps

Three simpler products, one per persona: dispatcher, QA, exec dashboard. Tempting: each role sees only what it needs, and each screen is clean in isolation.

Why I killed it. It re-created the exact disease I was curing. Three apps means three sources of truth and a reconciliation tax, which is what the spreadsheet already was. A coordinator cannot see why Processing is backing up without the QA queue behind it. I kept one shared model with five lenses.

03

AI auto-dispatch and full auto-triage, first

Let the system assign technicians and auto-resolve exceptions, minimizing the human from day one. Tempting: it is the cleanest leverage story, it demos beautifully, and “AI ops” gets a project funded.

Why I killed it. Ops leaders will not hand enterprise SLAs to a black box, and when an automated assignment blows a deadline, no one owns it. You also cannot build a routing model before instrumenting the manual decisions that would train it. I sequenced automation to v2 as suggestion-with-override.

04

A time-based Gantt as the primary board

The main surface as a horizontal timeline, one swimlane per job, bars against the SLA clock. Tempting: SLA is about time, and a Gantt shows time natively and looks planned.

Why I killed it. It optimized for planning, not real-time triage, and buried the 10% at-risk jobs inside a wall of on-track bars. I demoted the timeline to the job-detail view, where its real payoff lives: showing that most elapsed time was wait time between stages. Right idea, wrong altitude.

Information architecture

The system before the screens

For a data-dense tool, the object model is the design. I modeled the domain with eng before drawing a screen, so the UI would mirror the real database, not a convenient fiction.

Object model
OrderCaptureBundleProcessingJobQAReviewDelivery
TenantContract / SLAOrdersTechnicianAppointmentsCapacity
Pipeline state machine · discovered, not invented
OrderedScheduledCapturingUploadingProcessingQA PendingApprovedDelivered

QA Pending can loop back to Fix or Recapture. Any stage can raise an Exception, a first-class object carrying a defect code (BLUR, LOW_OVERLAP, COMPUTE_FAIL, MESH_HOLE, NO_SHOW). The state machine maps directly to the board columns and the card status token.

PulseOps · Command center for a 3D-capture network

Two flows where the network hands off to a human

PulseOps runs capture-to-delivery on autopilot and pulls a person in only for the ~10% of jobs that need judgment. These are the two moments it chooses to interrupt someone: a coordinator protecting an at-risk SLA, and a QA reviewer ruling on a 3D model.

Flow 01Clearing a P1 before the SLA clock runs outOps coordinator (Renata), managing by exception across the live pipeline
Wallboard flags a P1 at-risk jobOpen the deduped exception queueRead wait-time on the job timelineRecoverable inside the SLA window?No: log breach, alert account manager on client detailReassign via the Fleet capacity bandJob re-enters pipeline, clock re-baselinedException cleared, board back to green
A cleared exception drops the coordinator back to the queue; the next at-risk job surfaces itself and sorts to the top, so she reacts to interrupts instead of polling the whole network.
Flow 02QA verdict: accept, fix, or recapture a modelQA reviewer ruling on a flagged 3D model with keyboard decisions
QA Pending model routed to reviewerInspect the model in the 3D viewerClean enough to accept?Yes: keyboard-accept, Approved then DeliveredPin the defect, tag its codeFixable in compute, or recapture?BLUR / LOW_OVERLAP: recapture, job back to CapturingRoute MESH_HOLE fix to compute pipelineGPU reprocess, re-enter QA PendingApproved on re-review, delivered
A reprocessed model re-enters QA Pending and loops to step 1 for a second verdict; a recapture call loops the job all the way back to Capturing and a new field visit.
Design system for density

Color is a scarce resource I spent only on status

Density becomes legible only when it is systematic. One canonical taxonomy, color plus icon plus label so it never relies on color alone, used identically on the kanban card, QA queue, analytics table and client-health chip. A red token means the same thing everywhere, so a red token is a genuine alarm.

Healthyon-track · accepted · delivered
UploadingI/O in progress
Processingcompute · reconstruction
At-riskWIP high · SLA nearing
BreachSLA breach · rejected
Idlequeued · neutral
SLA at-risk
7
2 critical< 3h to due
18h 40m3h 30m0h 52mdelivered
Command Overview
Live wallboard
Updated 3s ago
⌘K
6
Network healthy·7 at-risk·+6.2% vs target6 open exceptions
Captures
38
12 in field now
In processing
24
WIP 20backing up
In QA
18
3 breachingwait 7.8h
Delivered
214
6.2%vs 200
SLA at-risk
7
2 critical<3h due
Fleet util
79%
in 75–80 band
Live pipelinethroughput today
Capture
12
Upload
9
1 at-risk
Processing
24
2 at-risk
QA
18
3 at-risk
Delivered
214
The Command Overview wallboard: the whole system composed into one glance layer, exceptions ranked on the right rail.
Sketch to ship

From a paper wall to a white workstation

The whole tool started as five dashboards taped to a wall, each carrying one open question in the margin. Every fidelity pass answered a few of them and threw away the parts that looked impressive but read as noise: a wall of paper first, then a stripped greyscale board, then a white workstation where color is spent only on status.

Lo-fi · paper
Locked the surface inventory and the manage-by-exception bet: five landing boards over one shared object model, with the open questions written in the margins before a single pixel was spent.
Mid-fi · greyscale
1Aging heat on left edge2At-risk sorts to top3WIP limit turns header amber
Killed the fourteen-field card and locked the hierarchy in pure greyscale, so nothing leaned on color: aging heat on the border, an at-risk-first sort, and WIP limits that turn a column header amber.
Hi-fi · shipped
Command Overview
Live wallboard
Updated 3s ago
⌘K
6
Network healthy·7 at-risk·+6.2% vs target6 open exceptions
Captures
38
12 in field now
In processing
24
WIP 20backing up
In QA
18
3 breachingwait 7.8h
Delivered
214
6.2%vs 200
SLA at-risk
7
2 critical<3h due
Fleet util
79%
in 75–80 band
Live pipelinethroughput today
Capture
12
Upload
9
1 at-risk
Processing
24
2 at-risk
QA
18
3 at-risk
Delivered
214
The shipped Command Overview: the whole system composed into one glance layer, exceptions ranked on the right rail.
Lo-fi · paper

Five boards, taped to a wall

Locked the surface inventory and the manage-by-exception bet: five landing boards over one shared object model, with the open questions written in the margins before a single pixel was spent.

Mid-fi · greyscale

One card, stripped to four fields

Killed the fourteen-field card and locked the hierarchy in pure greyscale, so nothing leaned on color: aging heat on the border, an at-risk-first sort, and WIP limits that turn a column header amber.

Hi-fi · shipped

Color spent only on status

Locked the white operator tool and the canonical status taxonomy of color plus icon plus label, with calm real-time on one shared clock tick. This is the version handed to engineering.

Marquee flow

Deep dive: the live pipeline board

It is 2pm. Twenty-four jobs are backing up in Processing and two are two hours from an enterprise breach. Three hard problems had to be solved on one surface.

  • Spotting a stall among hundreds of cards. Aging heat on the left border, an at-risk-first sort, and a WIP limit that turns a column header amber when it backs up. The eye is drawn to the problem, not asked to scan for it.
  • How much on a card before it is noise. Progressive disclosure: the card shows ID, client, SLA countdown and a time-in-stage meter. Everything else waits for the detail workspace. This is the fix from the fourteen-field board that tested badly.
  • Real-time that never makes the board jump. One shared clock tick drives every countdown, and the board never re-sorts under an active cursor. Exactly one live pulse; no ambient jitter.
Job IM-84213
Pipeline · Processing
Updated 3s ago
⌘K
6
PipelineProcessingIM-84213
IM-84213Redwood Realty · 1428 Camino Real, San JoseProcessing24h rush · 2h 14m to due
ReassignExpediteRe-queueOpen in QA
Lifecycle timeline wait time = 66% of elapsed
Ordered08:12·system
28m to schedule
Scheduled08:40·@dana
assigned Marcus R.
Captured10:20·Marcus R.
92 scan positions · 31% overlap
Uploaded10:42·system
6.4 GB in 22 min
Processing11:04·GPU-07
41 min · 1 COMPUTE_FAIL retry
QA Pending11:45·queue
waiting 1h 12m
Capture summary
Rooms8
Area4,200 sqft
Floors2
OccupancyOccupied
DeviceiPhone 15 Pro
Scan positions92
Coverage96%
Align residual1.2 cm
Assignment
MR
Marcus R.
San Jose zone · load 4/5
Reassign (capacity-aware)
Exceptions & QA history
COMPUTE_FAILretry 1 · resolved
MESH_HOLEflagged in QA · pending fix
Activity log
11:45queued for QA
11:44@marcus expedited
11:04processing started · GPU-07
10:42upload complete
Drill into the record workspace. The lifecycle timeline is the systems-thinking artifact: it shows most elapsed time was wait time between stages, not touch time.

The hours pool in Processing, so that stage earns its own surface for the reconstruction engineer: GPU queue depth, per-node health, failure and retry rate, and the quality signals that catch a silent regression before QA.

Compute Pipeline
GPU fleet · reconstruction
Updated 3s ago
⌘K
6
Queue depth
18
6 running
Throughput
42
jobs / hr
Failure rate
3.1%
last 24h
Retry rate
6%
auto-retry
Avg runtime
5.1h
bottleneck stage
GPU fleet · 8 nodesrunningidlefailed
GPU-01running
IM-84190
88% util
GPU-02running
IM-84176
76% util
GPU-03running
IM-84120
81% util
GPU-04retry 1
IM-84213
0% util
GPU-05running
IM-84077
84% util
GPU-06idle
no job
GPU-07running
IM-83998
96% util
GPU-08running
IM-84051
79% util
Reconstruction queue18 waiting
1IM-842406.4 GBwaiting 2m
2IM-842383.1 GBwaiting 4m
3IM-842365.8 GBwaiting 7m
4IM-842312.4 GBwaiting 9m
Reconstruction quality signalssurfaced upstream of QA
Alignment residual1.2 cm
within 2.0 cm tolerance
Mesh-hole rate4%
trending up 1.2 pts
Measurement confidence96%
stable
The Compute Pipeline: watching the bottleneck stage. A failed GPU node surfaces inline; reconstruction-quality signals (alignment residual, mesh-hole rate) push upstream of QA to stop bad models before they consume a reviewer.
The differentiator

Deep dive: the 3D-model QA review workflow

The reviewer inspects a spatial 3D artifact inside a 2D operational tool and decides accept, fix or recapture, at speed and with consistency. This is the surface that makes PulseOps domain-deep, unmistakably distinct from the order-lifecycle study.

QA Review
18 models in queue
Updated 3s ago
⌘K
6
IM-84213Redwood Realty · 1428 Camino Real, San Jose24h rush · 0h 52m to due
RawReconstructedFloorplanMeasure
master bath 8.4 ft · spec 8.0 (+0.4)
LivingKitchenBed 1Master BathBed 2Shower24.0 ft8.4 ft (+0.4)25
Three panes: the SLA-sorted queue, the model viewer with layer toggles and defect pins, and the inspector with a gated rubric and a keyboard-first decision bar. Accept stays locked until every check passes.

Three moves do the work. A fixed defect taxonomy makes reject reasons analyzable: every code feeds the exception chart. A keyboard-first flow protects review velocity. And the fix-versus-recapture call is made explicit, because it is the money decision: a light fix keeps the 24h SLA, while a recapture is roughly twice the cost and usually blows it. Then the loop closes: reject codes roll up into a per-technician capture-right-first-time score that feeds dispatch. The system learns from itself.

The business layer

Turning “turnaround feels slow” into a bottleneck you can point at

The analytics surface answers one question: where is the network breaking today, and who is affected. Every metric drills through to the underlying orders. The payoff is the bottleneck-attribution bar.

SLA & Client Health
Last 30 days
Updated 3s ago
⌘K
6
Last 30 daysAll regionsAll clientsAll typesnumbers marked target are design goals, not shipped results
SLA met
91%target 96%
Cycle time
31htarget 24h
Throughput
214per day
QA accept
82%target 85%
Exceptions
9%of jobs
Cycle-time distribution
0h24h SLA72h+
Arrivals vs completionsarrivalscompletions
May 1May 30
Exceptions by reason
BLUR
24%
LOW_OVERLAP
19%
COMPUTE_FAIL
16%
MESH_HOLE
14%
NO_SHOW
12%
MEASUREMENT_OOT
9%
ACCESS_FAIL
6%
Median time in stage · where the hours gowait time = ~66% of total cycle
5.1h
7.8h
CaptureUploadProcessingQADelivery
Enterprise client health91% global average hides the tail below
ClientVol 30dSLA attainmentCycleOpen excTrendHealth
Redwood Realty1,240
94%
26h288 · Healthy
Meridian Homes870
92%
28h381 · Healthy
Sentinel Insurance640
86%
34h763 · At-risk
Cascade Property Group410
78%
41h1147 · Critical
Wait time is about 66% of total cycle, pooling in Processing and QA. That one bar reframes the roadmap from 'make the UI faster' to 'attack the queue.' An 88% global SLA average hid Cascade Property Group breaching at 78%, proof that a global average is dangerous and per-tenant health had to exist. Every projected figure is tagged TARGET.

The dangerous 78% is one click from the average. Drill into Cascade and the story is specific and defensible: a 16-point SLA slide over seven weeks, an exec escalation, renewal at risk, and a recommendation pointing back at the Processing bottleneck. This is the surface an account manager walks into a QBR with.

Cascade Property Group
Enterprise · 24h rush tier
Updated 3s ago
⌘K
11
ClientsCascade Property Group
Health 47 · Criticalrenewal at risk
Account teamPMSO
SLA attainment
78%
contractual 95%
Volume 30d
410
of 500 commit
Median cycle
41h
target 24h
Open exceptions
11
3 breaching
SLA attainment · 7 weeks down 16 pts
7 wks ago95% contractual floornow · 78%
At-risk jobs4 of 11 open
IM-839982100 BroadwayProcessing2h 14m
IM-8404441 K StreetCapture3h 10m
IM-840519 Mission StQA1h 36m
IM-8412088 Pine AveUpload5h 40m
Escalations & account timeline
Exec escalation2d ago

12 orders missed the 24h SLA last week.

Renewal flagged at-risk5d ago

CSM logged churn risk ahead of Q3 renewal.

QBR scheduledin 6d

Root-cause review with the account.

Recommended

Attack the Processing bottleneck for Cascade first: it is 5.1h of the 41h cycle. Expedite the 3 breaching jobs today.

Per-tenant health: the account a global average was hiding. Every enterprise client holds a different contractual SLA, so health is tracked, and defended, per tenant.
The other half of the equation

Capacity is physical, so the tool refuses to over-drive it

You cannot just process faster: capture is a truck-roll inside a coverage radius. The fleet surface makes that constraint visible and pairs utilization with capture-right-first-time, so the tool never rewards pushing a technician into quality failures.

Fleet & Capacity
6 technicians · 3 zones
Updated 3s ago
⌘K
6
Network utilization
79%
target band 75–80
Captures / tech-day
4.6
target 5.5
On-time arrival
93%
last 7 days
No-show rate
4%
last 7 days
Right-first-time
86%
capture quality
Technician utilization 75–80% healthy band
Marcus R.San Jose
88%5/d
Priya S.Oakland
81%5/d
Dana P.San Jose
72%4/d
Leo M.Oakland
69%4/d
Sam T.Sacramento
61%3/d
Nina K.Sacramento
58%3/d
Marcus R. over-driven at 88%. Rebalance to protect capture quality.
Coverage · ~35 mi service radiiSacramento gap
OaklandSan JoseSacramento
Sacramento demand tomorrow22h needed · 16 avail
TechnicianZoneCaptures/dayOn-timeNo-showRight-first-time
Marcus R.San Jose5.096%2%84%
Priya S.Oakland4.894%3%88%
Dana P.San Jose4.492%5%89%
Sam T.Sacramento3.290%6%82%
Utilization is drawn with the 75–80% healthy band in view. Marcus is flagged over-driven at 88%; Sacramento shows a capacity gap against tomorrow's demand before it misses SLA.
Failures & recovery

Three things I got wrong

A case study with zero failures reads as a decorator who was never trusted with a hard problem. Each of these broke when it met reality, and each one forged a principle stated earlier on this page.

Iteration 01 · killed after testing
Ordered8
IM-84200
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84199
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84198
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Assigned6
IM-84193
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84192
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84191
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Accepted5
IM-84186
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84185
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84184
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Capturing12
IM-84179
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84178
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84177
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Uploaded9
IM-84172
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84171
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84170
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Processing24
IM-84165
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84164
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84163
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84162
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
QA18
IM-84158
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84157
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84156
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Rework4
IM-84151
Sentinel
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84150
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84149
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
Delivered214
IM-84144
Redwood
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84143
Cascade
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
IM-84142
Meridian
1428 Camino Real
4,200 sqft
2 fl
mrodriguez
San Jose
occ.
06/28
due 6/29
cov 96%
92 pos
GPU-07
note: retry
48s to locate the one breaching job
“I can't tell which one is actually on fire.”
Ops lead, first usability session
?
The Everything Board: 9 columns, 14 fields per card, everything the same weight. In testing it took the ops lead eleven seconds to find the one breaching job. Density is not information.

The eleven-second board

I built

A pipeline card showing everything I had, roughly 14 fields, because in a dense tool “more information” felt like the point.

It broke

In the first review I asked the ops lead to find the most at-risk job on screen. Eleven seconds and a scroll. Everything was the same visual weight, so nothing was salient. The board became wallpaper.

I learned

Density is not the enemy; undifferentiated density is. This is where “density with hierarchy” actually came from.

The fix

The card collapsed to ID, client, SLA countdown and a time-in-stage meter. Aging heat on the left border, an at-risk-first sort, and a WIP-limit header that turns amber when a stage backs up.

The wall of red

I built

Live updates where cards re-sorted the moment state changed, plus a toast for every state transition so nothing was missed.

It broke

The board moved under the cursor and caused misclicks. And the SLA-breach toast looked identical to the routine “delivered” toast, so within a day people muted all of them. I had built alert fatigue on purpose.

I learned

Real-time has to be calm, and unranked alerting is worse than none, because it teaches people to ignore the signal. This forged “calm real-time” and the polling-over-websockets tradeoff.

The fix

One shared clock tick; the board never re-sorts under an active cursor; toasts became a single risk-ranked triage queue. Severity is expressed by position, not interruption. A P1/P2/P3 model and dedup by root cause takes a wall of ~60 undifferentiated alerts down to ~6 ranked fires.

The defeated rubric

I built

A fully gated QA rubric: before a reviewer could decide, they scored ~12 criteria. The intent, consistency and analyzable data, was real.

It broke

It doubled review time, which is fatal on a 24h SLA, so reviewers slammed default values to reach the decision. The data I collected was slower AND garbage.

I learned

A rubric that fights the reviewer’s speed will lose. The analyzable signal has to ride along with the decision, not gate it. This is “exceptions surface themselves,” applied to human review.

The fix

The accept path collapsed to a keyboard-first decision (J/K to move, A accept, R reject). Only reject requires structure: exactly one defect code from a fixed taxonomy. That single code is the signal I needed, and it feeds the exception chart and the per-technician capture-right-first-time score.

Failure 02 · the fix
First alert modelMute all
COMPUTE_FAILIM-84213just now
COMPUTE_FAILIM-84210just now
COMPUTE_FAILIM-84207just now
NO_SHOWIM-84204just now
COMPUTE_FAILIM-84201just now
SLA_BREACHIM-84198just now
COMPUTE_FAILIM-84195just now
BLURIM-84192just now
COMPUTE_FAILIM-84189just now
Every state change fires a toast, so the rational response is to mute all of them, and the real breach hides in the noise.
Alert fatigue I built by accident, then designed out
The fix · quiet by default6 real fires
P1pages on-call
P2sits in the rail
P3rolls into a daily digest
dedup by root cause
20 × COMPUTE_FAIL1 GPU-node incident
Alert fatigue: an unranked board surfaces ~63 alerts, nearly all red, and trains operators to mute them within a day. A P1/P2/P3 severity model and dedup by root cause lands at ~6 real fires. This is why the exceptions rail shows six, not sixty-three.
Exception Triage
6 open · quiet by default
Updated 3s ago
⌘K
6
P1
2
pages on-call
P2
3
in the rail
P3
1
daily digest
38m
median resolve · target < 45m
severity: allcode: allowner: all 1 unowneddeduped by root cause · 20 COMPUTE_FAIL → 1
P1 · Pages on-call2
COMPUTE_FAIL
IM-84213Redwood Realty
2h 14m
M@marcus6m
Re-queue
SLA_BREACH_RISK
IM-83998Cascade Property
2h 14m
unassigned2m
Expedite
P2 · In the rail3
NO_SHOW
SAC 10:00Sentinel Insurance
D@dana14m
Reschedule
BLUR
IM-84109Redwood Realty
3h 30m
P@priya22m
Recapture
LOW_OVERLAP
IM-84077Meridian Homes
4h 02m
D@dana31m
Flag QA
P3 · Daily digest1
MESH_HOLE
IM-84051Cascade Property
P@priya38m
Route to fix
The Exception Triage queue: ranked P1/P2/P3, deduped by root cause, quiet by default, a named owner and one-click remediation on every row. Designed as the landing surface to drop time-to-triage from ~20 minutes to under two (target), the whole reason coordination stops scaling with volume.
Prototyping & validation

Tested against the operator, not my taste

An internal ops tool is one I cannot dogfood, so I pressure-tested it against the two people who own the outcome: recurring design reviews with the ops lead who lives in the pipeline, and think-aloud runs with a coordinator working real, anonymized jobs off the clickable prototype. That pass is done, and it moved both the design and the thesis. Then I scoped a moderated study for coordinators who have never seen PulseOps, protocol below, ready to run the day a build is stable.

Track 01 · Expert reviews + coordinator think-aloud
Done

Ten build reviews with the ops lead and three moderated think-aloud runs with a coordinator, all on live network state, not a happy-path demo. Every place they hesitated became a to-do. Three catches changed the design; one moment changed the thesis.

Friction

The coordinator read the SLA chip as time elapsed, not time left, and mis-ranked which job to touch first.

Change

Made every chip a countdown to breach with a 'to due' label and color ramp, so the number always means urgency.

Friction

Working the queue top-down, she reopened a job she had already handed off, with nothing marking it as taken.

Change

Gave each exception a named owner and a claimed state, so a triaged fire visibly leaves her queue.

Friction

She distrusted the board the way she distrusted the old sheet: no way to tell if a number was live or two hours stale.

Change

Added a visible freshness stamp on one shared clock tick, so 'is this current?' is answered before she has to ask.

The pivot

Watching the ops lead take eleven seconds to find the one breaching job on a fourteen-field board, I stopped designing a screen an operator scans and started designing one that surfaces the exception itself, the moment manage-by-exception stopped being a slogan and became the layout.

Track 02 · Moderated study, coordinator power user
Scoped

When a job is quietly aging in a queue no one is watching, does PulseOps put it in front of the coordinator before a client complaint does?

Recruit
Who
Ops coordinators or dispatchers who protect a delivery SLA in a live queue today: logistics, field-service, or capture ops.
How many
Five, the point where each added coordinator stops surfacing new triage breakdowns.
Mix
Three who run a manual tracker of spreadsheet plus Slack, two on a purpose-built console, to separate tool habit from triage skill.
Screen out
Anyone who has seen PulseOps or any InsideMaps internal tool, so first-shift reactions stay uncoached.
Format
60-minute remote moderated session on a seeded dataset, think-aloud, one SEQ ease question after each task.
Task script
01

You are taking over the board mid-shift. In your own words, what is about to be late, and what do you touch first?

Names the at-risk jobs and the true P1 within 2 minutes, no spreadsheet.
02

Cascade's order finished processing at 11am and no one has looked at it since. Find it before it breaches.

Opens the stalled QA-pending job while its countdown is still positive.
03

The queue shows a BLUR failure on a job due in three hours. Handle it.

Claims the exception and routes fix-or-recapture without leaving the queue.
04

Turnaround 'feels slow' this week. Show me where the network is actually stuck.

Reaches the bottleneck-attribution bar and names the Processing and QA wait, unaided.
05

The Cascade account manager asks whether their SLA is healthy. Answer it.

Drills past the global average to Cascade's 78% and states the renewal risk.
Success bar
Pre-registered · target
Time-to-triage (login to first correct action)< 2 min
Breach-spotting task success≥ 4 / 5 unaided
Breaches missed during the session0
Ease per task (SEQ)≥ 5.5 / 7
Overall usability (SUS)≥ 75
Collaboration & tradeoffs

How I worked, and what I chose not to build

I co-defined the status taxonomy with ops and eng so the UI states matched real database states, negotiated the refresh strategy against performance cost, and scoped the MVP with PM around the highest-leverage surface. The hard part of senior work is owning the downsides out loud.

Polling over websockets for v1

Accepted up-to-5s staleness to ship inside the eng budget. Mitigated with a visible freshness stamp and one shared clock tick, with a documented path to streaming.

One dense board over role-simplified views

Accepted a steeper onboarding curve to keep a single source of truth. Mitigated with saved filters and role-based default landing views.

Shipped the pipeline board and exception queue first

Sequenced Fleet, QA and analytics after. Cut anything not on the critical path to protecting SLA. The highest-leverage surface earned the first sprint.

Projected / target outcomes

The targets I designed toward, and how I would validate them

These are the metrics the system was designed to move, stated as targets, not shipped results. Each pairs a leading UX mechanism with the lagging operational outcome it drives, and each would be measured from stage-transition telemetry in a phased rollout that controls for confounders like seasonal volume.

On-time delivery within SLA
target 98% · projected
= the sum of four stage wait-times
1
Capture → Upload
surfaced by
Pipeline board
2
Upload → Processing
surfaced by
Job timeline
3
Processing → QA
surfaced by
Analytics bottleneck
4
QA → Delivery
surfaced by
QA queue
The legacy tool measured zero of these four leaves. You cannot manage a number you cannot see.
Design mechanism (leading)Operational outcome (lagging)
Time-to-triage15–20 min → under 2 min
SLA adherence88% → 96%+target

A single risk-ranked triage queue as the landing surface replaces reconstructing state from three tools each shift. Measured from login to first correct action.

Per-stage aging visibility0 → every stage
End-to-end cycle time (P90)72h → 40htarget

Wait time is where the hours are. Aging color-ramps plus WIP limits make dwell a routable signal and stop QA silently backing up. Tightening the P90 tail matters more than the median.

Manage-by-exception queuewatch all → touch ~10%
Ops headcount leverage+30–40% jobs / coordinatortarget

The system watches every job and surfaces only the exceptions. This is the metric that ties the design pattern to the business model: 2–3× volume at roughly flat coordination headcount.

Fixed QA defect taxonomyfree-text → 5 codes
First-pass QA yield70% → 85%target

A fixed taxonomy cuts reviewer variance, and surfacing capture-quality signals upstream catches poor captures before they consume processing.

Capture-quality feedback loopcaught at QA → at capture
Rework / recapture rate12% → under 7%target

Recapture is the most expensive failure. Pre-flight validation plus a QA-to-technician score shifts detection left, from QA days later to the capture moment.

Utilization band, cappeduncapped → 75–80%
Cost per delivered modelbounded, not maximizedtarget

Deliberately bounded and paired with first-pass yield, because over-driving past ~85% explodes variance. Guarding the ceiling is the defensible move (queueing theory).

Reflection

What I would validate first, and where PulseOps goes next

The first thing I would test with real usage: does per-stage aging actually reduce wait time, or merely make it visible? Visibility is necessary but not sufficient. If a coordinator can see a stall but lacks the authority or capacity to clear it, the design has surfaced a problem it cannot solve, and the next move is a workflow one, not a dashboard one.

The v2 I scoped: streaming to replace polling, predictive SLA-risk scoring calibrated hard against alert fatigue, suggestion-with-override auto-routing trained on the QA feedback loop, and an explicit, visible multi-tenant priority policy so ops can justify why one enterprise job jumps another. The honest hard part was resisting the urge to show everything. The first three board layouts were too dense to act on, and the version I handed to engineering is the one that hid the most without hiding anything that mattered.

Let's build something people remember

From enterprise teams to growing startups.

Let's talkarifin.yeasin@gmail.com