California's ALERTCalifornia network shows where production computer vision becomes an operational system. In 2025, it helped Cal Fire respond to roughly 3,600 wildfire incidents, with cameras detecting smoke before the first 911 call in more than half of them. The cameras accelerate detection; Cal Fire retains control over the response that follows.
Galicia runs a similar model. Its system, Xeocode, scans camera and satellite feeds with convolutional neural networks trained to catch a rising smoke column or heat signature. A technician confirms each match before crews deploy — the region built that check into its response after a record-breaking 2025 fire season, part of a €213 million prevention plan, its largest ever.
Neither system treats a correct detection as the finish line. That distinction matters because most computer vision investment still stops there — on accuracy, object classes, frame rate, the part a benchmark can score. What happens after a correct detection is harder to measure, and it's where production systems tend to fail: an alert reaches the wrong person, arrives outside the response window, or triggers an action that nobody verifies.
The detection itself is also less stable than the benchmarks suggest. Gartner predicts cross-modal vision models will surpass human-level scene understanding and predictive capability. A 2025 benchmark found deepfake detectors scoring 0.87–1.00 AUC on their original academic test sets falling as low as 0.43 against material from 2024. The published scores said nothing about how the models would hold up once the input distribution moved.
Newer architectures and bigger training runs keep pushing benchmark scores up. Production systems are now being asked to decide what should happen after a detection.
A forklift approaches a person at a blind corner. The system may detect both objects, interpret their movement, estimate the risk of a collision, or intervene. Those capabilities are often presented as one system even though each depends on different evidence.

Detection establishes what is present. Shelving, glare, and motion blur can degrade the image before the model has enough information to identify either object reliably. Model performance cannot recover details the sensor never captured.
Understanding establishes how the detections relate to one another. In this scene, “closing distance” depends on movement across consecutive frames. When one object blocks the other, the system may lose the original track or assign the returning detection to the wrong object.
Prediction estimates what is likely to happen next. Its input is a trajectory that may stop being valid before the response window closes, so the forecast has to update when either the person or the forklift changes speed or direction.
Prescription turns that forecast into an intervention, such as reducing the forklift’s speed. The response must arrive before the risk window closes, and a separate check must establish whether the intervention had the intended effect. Recording that a command was sent only confirms that the action pipeline ran.
Model choice is often the least consequential decision in a computer vision stack, even though it's usually the one debated first. Sensor choice and inference topology constrain the system before model evaluation begins. They determine what evidence reaches the pipeline and whether a response can arrive within the required operating window.
The input layer has the least room for error of any decision in the stack. Consider a warehouse safety system built on RGB cameras alone. Under normal conditions, it detects people and equipment reliably. In dense smoke, exactly the kind of condition the system exists to handle, the camera may still produce an image while capturing no usable evidence that the obscured people or equipment are there.
Noise degrades evidence the model can still weigh; smoke can remove the relevant visual evidence outright. Retraining or upgrading the model cannot recover information that never reached the sensor in usable form. A thermal sensor can preserve a signal where RGB imagery no longer supports a reliable decision.
If a camera-only installation later proves inadequate, adding another modality requires new hardware, mounting, cabling, and downtime. The sensor decision has already become an infrastructure decision.
Inference topology determines which stage the system can support reliably. A cloud-hosted prediction may still be useful when it informs a later decision. Prescription adds a hard deadline because the action has to arrive before the risk window closes.
The speed-reduction command must reach the forklift within two seconds. A cloud pipeline works only if transmission, queueing, inference, and actuation stay inside that budget, including latency spikes. A faster model cannot recover time lost elsewhere in the pipeline. Sensor choice fixes what evidence enters the system; inference topology determines whether the response arrives in time.
Mordor Intelligence identifies edge deployment as the fastest-growing segment of the computer vision market. That pattern fits what happens when cloud-hosted CV moves from analysis into control: the pipeline that was adequate for prediction can miss the deadline for action. Hybrid systems keep the time-critical trigger close to the device and leave slower analysis or model management in the cloud. Once the response window is fixed, inference location sets the system’s ceiling.
YOLO26 removes Non-Maximum Suppression, the post-processing stage used to discard overlapping detection boxes after inference. Its cost varies with the number of candidate boxes in a frame, so visually busy scenes can take longer to process even when the model itself runs at a stable speed. Dropping NMS removes one source of frame-to-frame latency variation. Queueing, preprocessing, hardware contention, and actuation can still introduce delays elsewhere in the pipeline.
In a prescriptive system, tail latency sets the safety margin. The forklift's speed-reduction command must still arrive within its two-second response window. A model that averages 20 ms but spikes to 200 ms is more dangerous in a closed-loop system than one that consistently runs at 35 ms, even if the first wins the benchmark chart.
Ultralytics reports up to 43% faster CPU inference for YOLO26 than YOLO11n. Its model-comparison page lists 38.9 ms versus 56.1 ms, which works out to roughly 31%, as LearnOpenCV's independent review also notes. The company expresses the same benchmark as two different-sounding numbers — a latency reduction and a throughput gain — without reconciling them for the reader, which is the more useful thing to know than either figure alone.
The point that carries forward is narrower than "YOLO26 is fast": removing a variable-cost step from the pipeline is what makes a model trustworthy enough to sit inside a closed-loop system in the first place.
Everything up to the action layer can work as designed and still leave the loop open. A trigger firing at the right moment confirms only that the detection path worked. The system must then verify that the action produced the intended result. The warehouse system triggers smoke suppression immediately and uses the same camera to verify the result. The camera is still looking through the same smoke, leaving the outcome of the intervention unclear. Verification becomes another vision task with its own failure modes. Reusing the same camera reproduces the trigger’s blind spot; a heat sensor can check whether the temperature actually falls.
BMW Group's Plant Regensburg keeps the final check human, and not by choice of caution. Deflectometry flags paint flaws, sanding robots correct them, and the line checks the result afterward. But the robots can't reach the body's edges, joints, or the fragile fuel filler flap — so a trained employee finishes those areas by hand, guided by a laser projector marking exactly where the flaws were recorded. At roughly 1,000 cars a day, the human check isn't a slower backup to an automated one; it's covering ground automation was never built to reach.
A closed loop requires an independent verification signal and a defined response when that check fails. The system may escalate or flag the mismatch for review. Without that step, the pipeline has acted without establishing whether the action worked. That level of automation may still suit a narrow, low-ambiguity task, provided it is described accurately.
Deployment scope depends on whether a wrong answer surfaces before it causes harm, through technical verification, human review, or a regulator’s paper trail, and what that safeguard costs to operate. Systems get scoped to that answer, not to the model’s ceiling.
Radiology accounts for most FDA AI/ML device authorizations, according to a 2025 systematic review in JAMA Network Open. But approval isn’t validation: most of those devices cleared through the FDA’s fastest pathway, built for incremental updates to existing device categories rather than new diagnostic judgment. In practice, clearance can rest on similarity to an existing device without independent clinical performance or safety data. Only 29% had any clinical testing at all.
Only 8% were tested with a human operator, despite being designed for use alongside one. “FDA-cleared” tells a hospital buyer the device resembles something already on the market. It says nothing about whether the device was ever tested with a clinician actually using it.
Viz LVO flags large-vessel occlusion strokes on CT angiography and sends simultaneous alerts to the emergency physician, radiologist, neurologist, and interventionalist. This removes the delay of sequential manual review. A cluster-randomized clinical trial across four stroke centers found that it cut door-to-intervention time by 11.2 minutes and CT-to-treatment time by 9.8 minutes. A 2025 systematic review and meta-analysis reported similar reductions in door-in-door-out time and time to specialist notification across pooled studies.
More than ten healthcare institutions adopted SciForce’s EfficientNet-B7 chest X-ray tool within six months. It flags signs of tuberculosis and COVID-19, but a generic classifier was too broad for triage: it marked almost any deviation as urgent, and a queue where everything looks urgent isn’t a queue at all. Simpler computer-vision methods had already failed on an earlier ECG project for the same reason: they missed subtle, atypical signs of disease.
We moved to EfficientNet-B7 and added layers to reduce false positives. Training labels came from radiologists tagging images directly and from NLP extracting labels from existing reports. Evaluation used unseen data from a public dataset and scans supplied by separate hospitals. A model that works only on one hospital's machines will not transfer reliably to the next.
A radiologist makes every diagnosis. The system handles the reversible part of the workflow by moving urgent-looking scans to the front of the queue. Critical cases were reviewed 30–40% faster, saving radiologists three to five hours per day.
Retail computer vision covers decisions with very different error costs. In pharmacy staffing, a missed queue alert usually means a longer wait. Checkout automation assigns products and charges to a transaction, so an incorrect result creates a billing error that compounds across thousands of transactions a day.
One industry consultant called computer-vision checkout "the hardest problem to solve" in retail. Amazon's Just Walk Out system combines cameras, 3D scans, a product-image catalog, and shelf-weight sensors to determine what a shopper picked up — genuine sensor fusion, not a single camera guessing. Walmart estimated that a comparable system would cost $10–15 million for a single 40,000-square-foot store and decided against it.
Amazon later withdrew the technology from its own large-format grocery stores and concentrated deployment on small and mid-sized third-party locations, where the economics were more favorable.
Across more than 1,000 pharmacy locations, SciForce’s staffing platform detects queues without following individual shoppers. Many systems calculate dwell and wait times by tracking people across frames. This one uses snapshot-based inference, so no customer is identified or followed.
Each frame is evaluated independently. Three or more people within 1–2 meters of the counter count as a queue. One ceiling-mounted wide-angle camera covers most stores, and alerts use channels the pharmacy already has, including a Telegram bot in some deployments. The alert itself requires no POS or CRM integration. Cross-referencing the camera feed with order records also distinguishes a quick online pickup from a longer in-person consultation.
The design sacrifices per-shopper wait-time estimates. It reduces the amount of video data that needs to be secured and avoids identity or demographic tracking. Lightweight computation allows the same modest hardware to work across stores with different layouts, ceiling heights, and lighting, with minimal site-specific tuning.
Staff response during peak periods improved by 40–60%, while identification of peak service hours improved by 30%.

Tracking a box across overlapping cameras becomes difficult once it has been flipped, stacked, or handed off between views. SciForce’s warehouse system addressed that re-identification problem for a high-volume beauty retailer, linking each box seen by one camera to the same box appearing later elsewhere in the facility.
Standard YOLO models needed additional training to distinguish packaging types and bags. Box size was estimated relative to nearby objects because camera angle made direct measurement unreliable.
Detection was fine-tuned on real warehouse footage. Large, high-contrast box IDs improved cross-camera readability, while barcode scans at the packing station independently confirmed that each physical box contained the expected product.
A box appearing in the wrong part of the warehouse triggers an alert for staff review. Human correction is economical here because a mistaken location alert is cheap to inspect. The camera feeds also generate heatmaps of worker movement, revealing bottlenecks between storage and packing stations that can be addressed through layout changes.
These workflow changes reduced misrouted shipments by 58% and improved order processing time by 19%.
A drone scan of a roof produces a dense, noisy mesh. Standard simplification can distort the sharp edges and corners that determine measurement accuracy. A 2024 study on 3D building reconstruction compared planar-surface-constrained decimation, designed to preserve roof faces and edges, with standard quadric-error simplification across nine real buildings. The generic method produced 25–31% more geometric error, with the largest differences on buildings containing the most distinct planes and edges.

A U.S.-based startup hired SciForce to turn drone footage into roof measurements for insurers, including area, length, slope, and joint type. Image capture was straightforward. The difficult part was converting an over-detailed 3D reconstruction into geometry precise enough to measure. The raw mesh contained thousands of tiny triangles, with trees sometimes represented at nearly the same polygon density as the roof. A rounded corner or a tree classified as part of the roofline could produce an incorrect figure on an insurance claim.
SciForce tested several decimation approaches, including standard quadric-error edge collapse and a custom planar-clustering method. The final process preserved roof edges and planes while reducing the mesh to a tenth of its original size or less.
A neural network trained on mesh geometry handled most roof-versus-background classification. Color analysis recovered distinct surfaces such as red or blue tile that geometry alone sometimes missed. Moss-covered roofs blending into vegetation required greater reliance on the geometric model.
The system produces the measurements an underwriter needs, with no downstream control action. Damage interpretation and claim valuation remain part of the underwriting workflow. Processing time fell from hours to minutes per property — an 80% reduction — with results matching field measurements within industry tolerance, at 99% measurement accuracy.
A threshold is usually selected on validation data by optimizing a metric such as F1 or accuracy. Unless the metric is cost-weighted, false positives and false negatives are counted without regard to their different production costs. That gap doesn't survive the handoff to production. The number ships, but the reasoning behind it does not. Infrastructure receives a threshold, not the cost model that produced it. Engineering wires a trigger to that number without knowing whether it was tuned for balanced test conditions or for the real costs of deployment. Months later, ops sees the threshold's real-world performance change, with no record of what it was meant to protect against.
This is why data science, infrastructure, engineering, and ops often disagree about what the system is actually optimizing for, even when everyone signed off on the same number. The number itself was never the problem — the missing cost model behind it was, and no one wrote it down.
The same handoff problem repeats at every layer below the threshold: a value or setting ships, and the reasoning behind it stays with whoever chose it.
A wrong threshold acts on noise or misses the failure the system was built to catch — most tuning starts before these are known:
With those in hand, compute:

This is your threshold — not 0.5. It's the cutoff where the expected cost of each kind of mistake is equal.

Two checks on the result, before you deploy it:
Google's What-If Tool will compute this for you if you'd rather not by hand — but check what it's defaulting to. Its cost ratio starts at 1.0 until you tell it otherwise, which is the same unexamined assumption most teams ship without noticing.
Before picking edge, cloud, or hybrid, get one number: the actual latency budget — the same distinction Architecture drew between prediction and prescription, applied here as a measurement rather than an argument.
Time the real window. From the moment the event occurs to the moment the response must land, in the actual deployment scenario — not a benchmark, not an assumption. If the trigger needs to act within two seconds of an event, that's your number. Nothing else below matters until you have it.
Compare it against each option's realistic round trip, not its best case: 2.1. A cloud round trip under normal network conditions, not a best-case ping. 2.2. An edge device's processing time under real, not lab, hardware load. 2.3. A hybrid split where only the time-critical piece runs local.
Pick the option that clears your budget with margin — not the one that clears it on average. A model that's fast most of the time and occasionally slow is more dangerous in a closed-loop system than one that's consistently a little slower.
If none of the three options clears the budget even with margin, the fix isn't a faster model — it's back to the action layer. Either the response deadline itself can be relaxed, or the trigger needs to fire earlier in the pipeline, before the event is this close to its deadline.
Why this is worth doing before anything else in the stack: reversing a hosting decision isn't a config change. It's new hardware, new mounts, a new procurement cycle. Of the four decisions here, this is among the most expensive to walk back — which makes it one of the most worth measuring before deciding, not after something's already running slow in production.
Before a trigger goes live, it's worth checking whether the action is actually built to hold up once it's live. Three questions worth asking:
Write each answer down before the trigger goes live, next to the config value it justifies. Without a documented answer to all three, the loop has a gap — one confident action with nothing behind it to catch a mistake.
The first three decisions get made once, at launch. This one runs for as long as the system stays live, and it's the one teams most often forget to schedule.
If the last calibration check happened before launch and never since, don't wait for a scheduled review to find out — assume drift is already underway and simply unmeasured, and run the check now. Document the cadence itself and the date each check last ran, not just the plan to run them — a schedule nobody can point to is the same as no schedule. Domain shift (the world in front of the camera no longer looks like the training data), dataset bias (the training data wasn't representative to begin with), and shortcut learning (the model latched onto a correlation that doesn't hold outside the training set) are the three failures usually hiding behind a drift alert — SciForce's breakdown of why computer vision models degrade after launch covers how each one shows up in production and what a monitoring pipeline actually needs to catch it.
Recognition was never the hard problem — most CV "breakthroughs" still solve it anyway, because it demos well. The harder problem doesn't show up on a benchmark: whether a system knows what to do once it's confident, and whether anything confirms that action worked. It surfaces months later, in production, whenever someone happens to notice.
The same gap opens up anywhere confidence gets treated as permission to act — agentic tool-calling, robotics turning detection into motion. Longer loops raise the stakes and shrink the evidence trail. Worth asking: at the moment your system is most confident, what confirms it's right? If nothing does, how would you find out before production does?
SciForce can help trace where a system is losing reliability before that gap shows up in a recall instead of a code review.