In 2025, Waymo's driverless vehicles crossed a threshold: independent, peer-reviewed data, not company demos, backed up the safety claims. A peer-reviewed analysis of 56.7 million rider-only miles found a 92% drop in pedestrian injury crashes compared to human-driver benchmarks, plus a 96% reduction in intersection crashes and an 82% reduction in cyclist and motorcyclist injury crashes. Waymo's own dashboard has since tracked the pedestrian figure at 220.6 million miles — more than triple the study's sample size — and it still holds at 93%.
The safety gain comes from sensor fusion. Waymo's Driver combines LiDAR, cameras, radar, and onboard compute into a continuous read of the scene before planning its next action. Each sensor contributes a different signal, so uncertainty in one stream can be resolved before it becomes a driving decision. That's multimodal AI architecture working as intended. The McDonald's IBM drive-thru pilot shows what happens when a system can't handle real-world ambiguity. It ran for years before McDonald's ended the test and removed it from more than 100 restaurants. In live lanes, the system made obvious mistakes, including adding nine sweet teas instead of one, because noisy and variable customer speech was harder to convert into accurate orders than the pilot could reliably handle.
Those failures became clear only after the system was already serving customers. This article covers the decisions that need to happen earlier: where speech and vision AI systems lose signal, why testing misses those failures, and how teams can catch them before they reach production.
Multimodal machine learning systems generally follow one of two architectural approaches. The first chains separate models together: one per modality, each converting its input into an intermediate representation before passing it forward — often text, but sometimes embeddings, structured features, or detected objects. A chained system that receives a text description of a prior scan can’t execute the instruction "compare with prior imaging" — the description tells it what was found, and spatial and temporal relationships between images often degrade after translation into text.
Combining speech and vision in AI systems through native multimodality means training the model from the start to process image, audio, and text in a shared representation, without translating each modality into a separate intermediate format before reasoning begins. Multimodal foundation models built this way include NVIDIA's Nemotron 3 Nano Omni, GPT-4o, and Gemini. native system can hold the encoded image, the note, and the prior scan in the same reasoning context simultaneously: the instruction "compare with prior imaging" becomes executable rather than unresolvable.

Large multimodal models process data in four broad stages: input, alignment, reasoning, and output. Knowing how large multimodal models work at each stage is important for diagnosis: a failure in the input layer and a failure in the reasoning layer can produce identical-looking wrong answers. Without knowing where in the pipeline the signal degraded, optimization effort lands in the wrong place.
Most multimodal accuracy gaps open in vision and audio encoders, before the reasoning layer processes anything, when training data stops matching production. Vision-language models such as SigLIP 2 and CLIP perform well on general visual tasks, but neither was trained on medical scans or factory inspection footage. Whisper handles dozens of accents and languages, but struggles with drive-thru noise or fast medical dictation.

The alignment layer causes more production failures than it gets credit for: in multimodal LLMs, each modality reaches it as a separate vector representation, and what gets projected into the shared reasoning space is often the actual point of failure.
Several leading open-source multimodal stacks use MLP projection — LLaVA-OneVision and Qwen3-VL both use it — because it is fast to deploy and cheap to run. It works by mapping visual tokens into the LLM's text space using a fixed transformation: every token processed the same way, regardless of what the image contains or what the task requires.
It works well for AI image understanding tasks with a clear visual answer: reading a label, parsing a chart, or identifying a product. Spatial reasoning is its weak spot. A fixed transformation has no way to treat a suspicious weld seam differently from the surrounding metal, or a lung nodule differently from surrounding tissue. When spatial task accuracy plateaus after launch, start with the alignment layer. MLP fails spatially by design — encoder problems tend to show up across task types, not specifically on location-dependent judgments.
A learnable Querying Transformer uses cross-attention to select which visual features reach the LLM, and that selection changes per input, per task, per region of the image. A chest X-ray processed by MLP sends every image patch forward with equal weight. Q-Former sends the ones that matter. It requires more compute and more training data than MLP, and that cost shows up at inference. The latency rules it out for high-volume, real-time pipelines — but in radiology and industrial inspection, where missing a localized detail isn't recoverable, it's the right architecture regardless of cost.
Both MLP projection and Q-Former resize and patch the image before alignment — the original aspect ratio is gone before the LLM sees anything. In surgical imaging, satellite imagery, and detailed industrial inspection, location is the finding. A 3mm nodule in the upper right lobe, a hairline fracture at a specific joint angle, an anomaly at a precise coordinate: resizing moves things. A model reasoning on a resized image is reasoning on a lie about where things are.
No production alignment technique solves this today. Treat it as an open architectural gap: budget for downstream verification on anything spatially precise until an alignment method that preserves native resolution reaches production.

After the alignment layer, all modality tokens enter the LLM backbone as a single unified sequence — image patches, audio frames, and text arriving as tokens, indistinguishable by origin. By this point, encoder and alignment problems should already be resolved. What reaches the backbone is assumed to be correct.
The backbone determines the reasoning ceiling. A multimodal system can read every input correctly and still draw the wrong conclusion, because input recognition and multimodal reasoning are two different jobs. The encoder did its job: each fact was captured accurately. Holding those facts together long enough to notice what they mean as a set is the backbone's job, and it's the one that fails first under context pressure.
A 235B parameter model running at full capacity at inference is not practical for most production systems. Mixture-of-Experts (MoE) solves this by activating only a fraction of parameters per token: the total parameter count and the actual inference cost are no longer the same number. Qwen3-VL has 235B parameters and uses 22B per token. NVIDIA Nemotron 3 Nano Omni pushes the same ratio further, at 30B total and 3B active. That gap is what closes the internal argument that production-scale multimodal inference is too expensive to run.
In production AI systems, the output layer is where model behavior becomes system behavior. A weak output design may not fail during testing because the downstream system has not yet received the edge case that exposes it. In production, the same model answer can be harmless as text, risky as structured data, and unacceptable as an autonomous action.
The most common enterprise output type in 2026, mainly because it requires the least downstream integration. In high-volume deployments, text output quality degrades under context length pressure. As the accumulated token sequence approaches the context window limit, the model begins to drop or compress earlier context. That window is an active engineering problem at the pipeline design stage.
When voice is the output, latency is the production constraint. General-purpose TTS pipelines typically introduce 300–500ms of additional latency at the output stage — well outside the 200ms window conversational turn-taking expects. NVIDIA Nemotron Speech addresses this by delivering production TTS as a NIM microservice built for low latency. Above the 200ms threshold, the silence feels unnatural and the speaker fills it with a repeated or rephrased input, triggering a second inference cycle on a modified signal.
The LLM generates JSON or structured records that feed directly into downstream systems like POS terminals, EHR systems, and quality control dashboards. Schema validation catches format errors, not factual ones: a hallucinated value that matches the expected format passes validation and enters the downstream system as confirmed. Constrained decoding fixes this at the source: a field limited to a predefined value list cannot produce a value outside it, because the model is never given the option to generate one
Agentic output skips the step where a human or downstream system reviews what the model produced — the system decides what to do next and does it. Unlike a wrong value in a structured data field that produces one wrong action, agentic decisions compound: each agent treats the previous agent's output as correct, so a single wrong decision moves through the system before anything looks wrong in isolation. Audit trails at the decision level and human-in-the-loop gates for high-risk actions are what separate a recoverable mistake from one that isn't.
Adding a second modality moves a multimodal system from detection to decision: from flagging that something needs attention to determining what to do about it.
Imaging-only AI has a consistent limitation: the model sees the scan but not the patient. A chest X-ray read in isolation misses smoking history, prior scans, lab values, and the clinical note that says the patient presented with shortness of breath three days ago. Without that context, a nodule is just a nodule. With it, the system can distinguish low probability of malignancy given a stable presentation from high priority for biopsy given progression and new symptoms.
SciForce's automated lung pathology detection system applies the same principle to radiology triage: the image model does not operate as a standalone classifier, and the clinical report is not treated as after-the-fact metadata. EfficientNet-B7 extracts features from chest X-rays, NLP pipelines process medical reports for clinical context, and a prioritization layer ranks critical cases for radiologist review. The system was trained and validated on ChestX-ray14 and client-provided records to make the pipeline less dependent on a single imaging source or patient population. Critical case review time dropped by 30–40%, with the architecture taking over the detection-and-ranking step so radiologists could spend more time on interpretation.
A multimodal ordering system knows when what it heard is inconsistent with what the menu allows, and routes to clarification instead of the POS terminal. The InTouch Insight 2025 mystery shopping study shows what's missing when that check isn't there: employee-assisted AI achieved 90% order accuracy versus 83% for AI-only, a 7-point gap produced by edge cases the ASR system couldn't resolve without contextual disambiguation. The system didn't know it was wrong because its confidence score reflected acoustic clarity instead of semantic accuracy.
SciForce's voice-driven ordering system shows what that disambiguation layer looks like in production. A custom Voice Activity Detection model separates customer speech from passenger chatter and engine noise without requiring a wake word, while the ASR engine runs on CPU hardware across hundreds of lanes, staying responsive enough for natural, real-time conversation. When confidence drops below threshold on a critical menu item, the system triggers a clarification prompt, like "Did you mean cheeseburger or fishburger?", instead of passing a low-confidence transcript to the POS terminal. Word error rate is tracked separately for core menu terms rather than as a single aggregate score, because a general WER can look acceptable while "combo," "fries," or another revenue-critical item is still misrecognized often enough to matter.

Average order time dropped from roughly 110 seconds to under 90 seconds per customer — an 18–25% improvement driven by fewer clarification loops and a streamlined handoff straight to the kitchen POS. The system's upselling logic (recommending items based on what's already in the order) adds revenue the original ASR pipeline never touched.
An AWS case study from December 2025 on Amazon fulfillment center predictive maintenance shows how cross-modal fusion works in this setting: a chatbot classifies vibration data against ISO 20816-1 thresholds, lets technicians upload photos, audio notes, or equipment video, and retrieves the matching procedure from repair manuals.
In SciForce’s work, cross-modal fusion uses separate encoders and late fusion when vibration and acoustic signals behave too differently to combine early without degrading both. The resulting accuracy loss is hard to trace: it shows up in overall accuracy, not in either modality's output individually. The 50% Undetermined rate is what that misstep looks like at scale. In multimodal AI deployment at the edge, fusion strategy becomes a deployment constraint rather than a model-choice question. Our 2026 predictive maintenance guide explains how edge conditions change the choice between on-device and offloaded inference, and why that choice affects both latency and monitoring accuracy.
A wrong pre-launch decision in a multimodal system is harder to find than in a single-modality one. Each additional modality adds a layer that can mask the problem, and by the time it surfaces in output, the rest of the pipeline has been built around it.
For regulated industries, multimodal AI infrastructure challenges begin with data protection: HIPAA and similar requirements often make on-premises or private-cloud deployment the default.
For regulated industries, compliance settles the multimodal AI architecture question before capability enters the conversation: HIPAA and similar data-protection regulations make on-prem or private cloud the default. For everyone else, multimodal AI infrastructure challenges increasingly matter more than model access: the Stanford HAI 2025 AI Index found that open-weight models narrowed the performance gap with closed models from 8% to 1.7% on the Chatbot Arena Leaderboard in a single year.
When degradation clusters around a specific input type, the underlying cause is a distribution gap. Fine-tuning on domain-specific data closes it: consistent failure on images from one scanner manufacturer, or audio from one acoustic environment, is the signal that points there.
A vision encoder fine-tuned on DICOM images outperforms a general ViT on medical imaging tasks, because fine-tuning changes what the encoder represents: prompt engineering only changes how the model is asked, and it can't compensate for a representation gap at the encoder layer. The same principle applies across vision and speech AI systems: fine-tuning the alignment projection on domain-paired data closes a cross-modal gap that no amount of backbone capability can substitute for. In practice, that's image-text pairs from your domain, or audio-transcript pairs from your environment.
Encoder fine-tuning requires hundreds to thousands of labeled domain samples for vision, hours of domain audio for ASR. Alignment layer fine-tuning requires paired cross-modal data. Backbone fine-tuning requires the most data and computation and is rarely justified unless general LLM reasoning doesn't transfer to the task.
Testing tells you whether the system works on known inputs. The failures below appear in production, on inputs no controlled environment generates, and by the time they surface in output the architectural window to fix them cheaply has already closed
Every modality needs a null handler: route to single-modality inference if one input drops, flag for human review if two drop, withhold the output entirely if the remaining signal falls below a minimum confidence threshold. When all modalities are present but contradicting each other, log each modality's output independently at the pre-fusion stage, define a divergence threshold per modality pair, and route outputs that exceed it for human review. Once fusion runs, the individual signals are gone: the conflict can't be reconstructed from the final output.

False agreement is harder to catch than conflicting signals. When two modalities confirm each other's wrong output, the system treats the agreement as evidence of correctness with no mechanism to recognize it as circular. A fraud detection model flags a transaction based on spending anomalies. A second model finds no prior fraud at the merchant, and the system uses that confirmation to approve the transaction. The merchant data was clean because the fraud ring had been operating there for less than 48 hours: the second model confirmed the absence of a signal that didn't yet exist, not the legitimacy of the transaction. Provenance tracking at the fusion layer catches this by logging the call sequence and identifying when one modality's output is downstream of another rather than independent.
The value of multimodal AI is architectural: a system with more than one way to reach a correct output keeps producing correct outputs when conditions change. That property has to be designed before the first inference runs.
If one modality in your system produced a high-confidence wrong output right now, would your pipeline detect it, or would it propagate silently until it arrived at the output layer looking correct?
SciForce builds production AI vision and speech systems across healthcare, HoReCa, industrial, and enterprise multimodal AI environments. If you're designing a multimodal system and want to pressure-test the architecture before the first inference runs, get in touch.