When an Air Canada customer asked the airline’s website chatbot about bereavement fares, it told him he could book first and claim the discount within 90 days. The same answer linked to a policy page saying retroactive requests weren’t allowed.
The passenger followed the chatbot’s instructions and later had his refund request rejected. The civil tribunal found the airline liable for negligent misrepresentation after concluding that he’d reasonably relied on the inaccurate guidance. The correct information existed on the same website, but the tribunal held that customers shouldn’t have to reconcile contradictory information across different parts of it.
Air Canada’s chatbot failure shows what can happen when an application produces an answer without reliably using or checking authoritative information. The case doesn’t reveal whether the chatbot used RAG, fine-tuning, or another architecture, but it illustrates the questions that needs to be answered when designing such systems: how information reaches the application, whether the model’s behavior needs training, whether the task belongs in generation or a structured system, and how the result is verified. Menlo Ventures’ 2024 enterprise AI survey among 600 US enterprise IT decision-makers found that production RAG use rose from 31% to 51% year over year, while fine-tuning remained at 9%.
Production systems fail when the wrong component gets the job. This article shows when a workflow needs RAG, fine-tuning, both, or neither, and how to recognize a bad fit from its failure patterns.
Before choosing between RAG and fine-tuning, teams need to separate three questions: What should the model be able to do? Where should it get the information it needs? And does the task require generation at all?
Model choice determines the capabilities available before application-specific data enters the workflow. Knowledge access determines how current, proprietary, or transactional information reaches the application at runtime.
A general-purpose model is usually the practical starting point for varied prompts and open-ended analysis. A domain-pretrained or otherwise adapted model may handle specialist terminology and recurring domain patterns more consistently, although specialization doesn’t guarantee better overall performance than a strong general model.
Model size is a separate decision. Smaller models may be sufficient for constrained workloads, but they still have to meet the application’s quality, reliability, and service requirements.
Information may come from model parameters, retrieved documents, databases, APIs, or business rules. RAG supplies document passages at runtime and can work with a general-purpose, domain-adapted, or fine-tuned generator.
Document retrieval suits policies, research, and other material that changes over time. Exact records, calculations, and transactional rules are usually better handled through structured systems than through retrieval or free-text generation.
Choosing the right information source still doesn’t guarantee a correct response. Before an answer reaches the user, the application may need to verify source support, enforce business rules, and route uncertain results to refusal, review, or a safer execution path.
Fine-tuning updates model parameters using examples from the target task, improving classification, extraction, terminology, output format, tool selection, and repeatability across similar requests. The training data may also introduce factual associations, but those associations are difficult to update selectively and don’t point directly to an authoritative source.
The approach works best when the required behavior is stable, measurable, and repeated often enough to justify training and maintenance. Changing or attributable information is usually better supplied at runtime through retrieval, databases, or APIs. Retrieval and adaptation can be combined when a workflow needs both current information and trained behavior. Broader systems may also route requests among generation, retrieval, and structured processing.
Some tasks don’t need a generated answer. Exact lookups, calculations, eligibility rules, account balances, inventory checks, and other transactional operations are usually better handled by databases, APIs, or deterministic logic. An LLM may explain the result, but it shouldn’t determine the underlying value. Fixed templates and narrowly defined classification or extraction tasks may also be better suited to conventional ML. When the possible outputs are known, these systems can be easier to evaluate and control than open-ended generation. The application still needs a defined response to missing records, invalid inputs, or uncertain classifications. Depending on the workflow, it may ask for clarification, use another processing path, send the result for review, or refuse the request.
SciForce developed a contract-processing system for a Swedish law firm. The system had to classify clauses and extract facts from partnership agreements, NDAs, and service contracts. Because Swedish-language NLP resources were limited, building a large language corpus or relying on translation would have been expensive and difficult to maintain.
The system combined Gaussian-process models with TF-IDF and PCA, as well as Word2Vec features and classification-tree ensembles. It identified clauses, entities and their relationships, references to arbitration and courts, and different types of disputes.
This case shows that legal text processing doesn’t always require generative AI. When the expected outputs are clearly defined and easy to measure, conventional ML can classify text and extract information directly.
A RAG system can fail before the model generates a single word. The corpus may contain an outdated policy, the retriever may miss the decisive passage, or context selection may bury it among weaker matches. Even when the right evidence reaches the model, it can still misread it, ignore it, or make a claim the source doesn’t support.
A retrieval pipeline can return a plausible passage and still fail the task. The problem may begin with the corpus itself:
Dense retrieval ranks passages by semantic similarity, but similar language doesn’t always mean relevant evidence. A 2025 legal embedding benchmark found that models leading general-purpose rankings didn’t retain the same positions across expert-annotated legal datasets. Retrieval quality must therefore be evaluated on the intended corpus.
The first pass must include the evidence required to answer the question. Re-ranking can improve its position, but it can’t recover a passage the retriever missed. Context selection then removes duplicates and weaker matches before the remaining evidence reaches the generator.
Finding the right evidence is only half the job. The model still has to use it correctly. Two things can go wrong:
Production checks therefore can’t stop at whether the system retrieved relevant text. Before releasing an answer, the application must verify that it answered the question, supported its claims with the cited sources, and reached the correct conclusion.

RAG is a strong fit when information:
This makes RAG useful for policies, research, internal documentation, and other evolving collections. But keeping information outside the model creates an operational responsibility. Documents must be added, removed, and re-indexed; permissions must remain synchronized; and retrieval quality must be checked as the collection changes. A stale document in a frequently refreshed index is still a stale document.
The application also needs a path for questions the retrieved material can’t support. It may retry the search, use a structured lookup, request human review, or refuse the answer.
A SciForce healthcare platform needed to map free-text clinical input, including shorthand notes and symptom descriptions, to SNOMED CT, LOINC, and RxNorm concepts. Fixed rules couldn’t reliably handle variations in clinical phrasing.
The system used GPT-based query normalization and a locally deployed Qdrant index to retrieve matching concepts from controlled terminology collections. The generator wasn’t fine-tuned.
Because clinical terminologies change, the system tracks new releases, removes deprecated concepts, refreshes the index, and sends updated mappings for expert review. The knowledge remains in a layer that can be updated and inspected without retraining the model.
Fine-tuning is worth considering when prompting can’t produce sufficiently reliable behavior for a stable, repeated task. The decision depends on what needs to change, whether the improvement can be demonstrated, and whether it justifies the full training lifecycle.
Fine-tuning can improve classification, extraction, terminology, output structure, tool selection, and consistency. It’s less suitable for changing information or claims that must link to an original source.

These options aren’t mutually exclusive. A model may undergo domain-adaptive pretraining followed by supervised fine-tuning. For LLMs, supervised fine-tuning commonly uses parameter-efficient methods such as LoRA or QLoRA to reduce the computing and memory requirements of adaptation.
A SciForce language-learning platform used fine-tuning to generate structured lessons, personalize exercises, and automate assessments while keeping outputs aligned with curriculum standards.
User profiles and course metadata were stored in application databases. Quiz results, completed activities, and recurring mistakes informed changes to exercise difficulty and future recommendations.
Educators reviewed the relevance of generated lessons, quizzes, and feedback, while students tested personalized exercises and writing and speaking assessments. The team used their responses and performance data to refine generated content and adaptive behavior before a gradual rollout. Continued updates addressed curriculum changes, educational standards, and new user feedback.
A completed training run isn’t evidence of improvement. Evaluation should pass three checks:
A 2025 cardiology study shows why evaluation scope matters. Researchers fine-tuned Llama 3.1-8B on real and synthetic discharge summaries, but three cardiologists reviewed only ten selected outputs. The results support feasibility for a defined documentation task, not reliability across institutions or uncommon cases.
The same tests should run after changes to the base model, adapter, dataset, prompt, or serving configuration.
Fine-tuning becomes an economic decision when its recurring value can be compared with its lifecycle cost:

Value per response may come from shorter prompts, fewer retries, less human correction, or lower latency. Useful lifetime is how long the trained behavior remains valid before labels, requirements, or the base model change.
A high-volume task may still be a poor candidate if it changes every few weeks. A lower-volume task may justify training when each failure requires expensive expert review.
A single application may receive requests for exact records, current documents, generated summaries, and repeated domain-specific outputs. Its architecture determines which component handles each request and what happens when the result fails validation.

Start with the type of result and its authoritative source. Then decide whether the model needs trained behavior. Define validation and fallback before comparing model sizes, since those execution paths affect both quality requirements and total cost.
These patterns can be combined within one execution path. For example, retrieval may supply current evidence to an adapted generator. A routed workflow operates at a broader level, selecting among generation, retrieval, structured processing, or a sequence of them according to the request.
Routing matters when one interface accepts several types of requests. “How many employees joined this quarter?” calls for a direct lookup. “Why did employee turnover change this quarter?” may require records from several systems followed by analysis and a written explanation.
Sending both requests through the same LLM adds unnecessary cost to the first and gives the second no reliable way to access the required data. A router can classify the request, check the user’s permissions, and send it to structured lookup, retrieval, generation, or a sequence of these components.
In SciForce’s enterprise data-processing case, the application combined information from HR platforms, CRMs, financial tools, and operational databases. Direct lookups handled requests such as employee statistics and sales figures. Vector search and client-specific knowledge bases supplied business information for source-dependent requests. The LLM was reserved for summaries, trend analysis, and other tasks that required synthesis.
Query-routing logic selected the execution path. Role-based permissions restricted access to client data, while query filters and response validation checked requests and outputs. Simple requests could return data directly; analytical requests could retrieve the required information before passing it to the model.
Model size should be selected after the workflow has assigned work to retrieval, adaptation, structured processing, and generation. The remaining question is how much model capacity that work requires.

A smaller model may perform poorly because it doesn’t have the information it needs, not because it’s too small. Before switching to a larger model, test both options with the same retrieval system, tools, prompts, and checks. If the smaller model meets the quality target without too many retries or fallbacks, a larger model may increase costs without improving the final result enough to matter.
AWS Bedrock assigns two Custom Model Units to Llama 3.1 8B with a 128K context window and eight to the 70B version. This gives the smaller model a clear capacity advantage on that platform.
That advantage depends on how many responses pass validation. Retries, calls to a larger fallback model, and human review increase the cost of each usable result. Test both models on the same production sample and calculate:

For a local deployment, replace the AWS capacity figures with the required GPUs, peak-traffic replicas, and ongoing maintenance. The cheaper option is the model that delivers an accepted response at the lower total cost.
Suppose a user gets a wrong answer tomorrow. Could you tell within an hour what happened? The system may have retrieved an old document, ignored the right passage, used the wrong model behavior, or sent a simple lookup through generation. You’d also need to know what should happen next: retry the request, check a database, ask for review, or refuse to answer.
Take the request your application can least afford to get wrong and follow it from input to final response. Check where the information comes from, which component handles it, and what catches a failure before the answer reaches the user.
If that path is hard to trace, SciForce can help you review and redesign it.