• Services
    LLM
    AI & ML
    Digital Healthcare
    Data Science
    DevOps
  • Products
    Jackalope
    EyeAI
  • Industries
    Healthcare
    Agriculture
    EdTech / LMS
    Retail / E-commerce
    Manufacturing
  • Resources
    Blog
    Case Studies
    Expert Guides
  • Company
    About us
    Careers
  • Contact us
logo
Services
LLMAI & MLDigital HealthcareData ScienceDevOps
Industries
HealthcareAgricultureEdTech / LMSRetail / E-commerceManufacturing
Case StudiesAbout UsBlogCareers
Our contacts
+380(66)54-32-579
sales@sciforce.tech

Get monthly digest of innovations

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Social Media:
Privacy Policy © 2026 Sciforce
5.0
RAG vs Fine-Tuning: How to Choose the Right AI Architecture

Building Domain-Specific LLM Systems: When to Use RAG, Fine-Tuning, or Neither

Published: September 18, 2026
# AI / ML
# Data Science
# LLM

Introduction

When an Air Canada customer asked the airline’s website chatbot about bereavement fares, it told him he could book first and claim the discount within 90 days. The same answer linked to a policy page saying retroactive requests weren’t allowed.

The passenger followed the chatbot’s instructions and later had his refund request rejected. The civil tribunal found the airline liable for negligent misrepresentation after concluding that he’d reasonably relied on the inaccurate guidance. The correct information existed on the same website, but the tribunal held that customers shouldn’t have to reconcile contradictory information across different parts of it.

Air Canada’s chatbot failure shows what can happen when an application produces an answer without reliably using or checking authoritative information. The case doesn’t reveal whether the chatbot used RAG, fine-tuning, or another architecture, but it illustrates the questions that needs to be answered when designing such systems: how information reaches the application, whether the model’s behavior needs training, whether the task belongs in generation or a structured system, and how the result is verified. Menlo Ventures’ 2024 enterprise AI survey among 600 US enterprise IT decision-makers found that production RAG use rose from 31% to 51% year over year, while fine-tuning remained at 9%.

Production systems fail when the wrong component gets the job. This article shows when a workflow needs RAG, fine-tuning, both, or neither, and how to recognize a bad fit from its failure patterns.

RAG, Fine-Tuning, and Structured Systems Solve Different Problems

Before choosing between RAG and fine-tuning, teams need to separate three questions: What should the model be able to do? Where should it get the information it needs? And does the task require generation at all?

Separate Model Capability from Knowledge Access

Model choice determines the capabilities available before application-specific data enters the workflow. Knowledge access determines how current, proprietary, or transactional information reaches the application at runtime.

  • Model Choice

A general-purpose model is usually the practical starting point for varied prompts and open-ended analysis. A domain-pretrained or otherwise adapted model may handle specialist terminology and recurring domain patterns more consistently, although specialization doesn’t guarantee better overall performance than a strong general model.

Model size is a separate decision. Smaller models may be sufficient for constrained workloads, but they still have to meet the application’s quality, reliability, and service requirements.

  • Knowledge Access

Information may come from model parameters, retrieved documents, databases, APIs, or business rules. RAG supplies document passages at runtime and can work with a general-purpose, domain-adapted, or fine-tuned generator.

Document retrieval suits policies, research, and other material that changes over time. Exact records, calculations, and transactional rules are usually better handled through structured systems than through retrieval or free-text generation.

Choosing the right information source still doesn’t guarantee a correct response. Before an answer reaches the user, the application may need to verify source support, enforce business rules, and route uncertain results to refusal, review, or a safer execution path.

Decide Whether the Model’s Behavior Needs Training

Fine-tuning updates model parameters using examples from the target task, improving classification, extraction, terminology, output format, tool selection, and repeatability across similar requests. The training data may also introduce factual associations, but those associations are difficult to update selectively and don’t point directly to an authoritative source.

The approach works best when the required behavior is stable, measurable, and repeated often enough to justify training and maintenance. Changing or attributable information is usually better supplied at runtime through retrieval, databases, or APIs. Retrieval and adaptation can be combined when a workflow needs both current information and trained behavior. Broader systems may also route requests among generation, retrieval, and structured processing.

When the Task Should Bypass Generation

Some tasks don’t need a generated answer. Exact lookups, calculations, eligibility rules, account balances, inventory checks, and other transactional operations are usually better handled by databases, APIs, or deterministic logic. An LLM may explain the result, but it shouldn’t determine the underlying value. Fixed templates and narrowly defined classification or extraction tasks may also be better suited to conventional ML. When the possible outputs are known, these systems can be easier to evaluate and control than open-ended generation. The application still needs a defined response to missing records, invalid inputs, or uncertain classifications. Depending on the workflow, it may ask for clarification, use another processing path, send the result for review, or refuse the request.

Legal Contract Processing Without Generation

SciForce developed a contract-processing system for a Swedish law firm. The system had to classify clauses and extract facts from partnership agreements, NDAs, and service contracts. Because Swedish-language NLP resources were limited, building a large language corpus or relying on translation would have been expensive and difficult to maintain.

The system combined Gaussian-process models with TF-IDF and PCA, as well as Word2Vec features and classification-tree ensembles. It identified clauses, entities and their relationships, references to arbitration and courts, and different types of disputes.

This case shows that legal text processing doesn’t always require generative AI. When the expected outputs are clearly defined and easy to measure, conventional ML can classify text and extract information directly.

How RAG Works in Production

A RAG system can fail before the model generates a single word. The corpus may contain an outdated policy, the retriever may miss the decisive passage, or context selection may bury it among weaker matches. Even when the right evidence reaches the model, it can still misread it, ignore it, or make a claim the source doesn’t support.

Retrieval, Re-Ranking, and Context Selection

A retrieval pipeline can return a plausible passage and still fail the task. The problem may begin with the corpus itself:

  • Source authority: a highly similar document may not be the source the application should trust.
  • Versioning and refresh: an obsolete policy may remain indexed while its replacement is missing.
  • Permissions: the retriever may surface material the user isn’t authorized to access.
  • Chunking and deletion: decisive context may be split across passages, duplicated, or retained after the source document should’ve been removed.

Dense retrieval ranks passages by semantic similarity, but similar language doesn’t always mean relevant evidence. A 2025 legal embedding benchmark found that models leading general-purpose rankings didn’t retain the same positions across expert-annotated legal datasets. Retrieval quality must therefore be evaluated on the intended corpus.

The first pass must include the evidence required to answer the question. Re-ranking can improve its position, but it can’t recover a passage the retriever missed. Context selection then removes duplicates and weaker matches before the remaining evidence reaches the generator.

The Right Documents Can Still Lead to the Wrong Answer

Finding the right evidence is only half the job. The model still has to use it correctly. Two things can go wrong:

  • The model ignores the source. In a study of medical question answering and HotpotQA, the correct evidence was already available, but models relied on information learned during training. This happened four to seven times more often than the retriever returning the wrong passage.
  • The model can’t connect the sources. A 2026 study of 11 models supplied every passage needed to answer scientific questions that required several reasoning steps. The models still produced the wrong answer in 58.6% of cases.

Production checks therefore can’t stop at whether the system retrieved relevant text. Before releasing an answer, the application must verify that it answered the question, supported its claims with the cited sources, and reached the correct conclusion.

The Right Documents Can Still Lead to the Wrong Answer

When RAG Is the Right Fit

RAG is a strong fit when information:

  • changes more often than the model should be retrained;
  • must be linked to an identifiable source;
  • is available only to particular users or roles.

This makes RAG useful for policies, research, internal documentation, and other evolving collections. But keeping information outside the model creates an operational responsibility. Documents must be added, removed, and re-indexed; permissions must remain synchronized; and retrieval quality must be checked as the collection changes. A stale document in a frequently refreshed index is still a stale document.

The application also needs a path for questions the retrieved material can’t support. It may retry the search, use a structured lookup, request human review, or refuse the answer.

Clinical Terminology Retrieval Without Fine-Tuning

A SciForce healthcare platform needed to map free-text clinical input, including shorthand notes and symptom descriptions, to SNOMED CT, LOINC, and RxNorm concepts. Fixed rules couldn’t reliably handle variations in clinical phrasing.

The system used GPT-based query normalization and a locally deployed Qdrant index to retrieve matching concepts from controlled terminology collections. The generator wasn’t fine-tuned.

Because clinical terminologies change, the system tracks new releases, removes deprecated concepts, refreshes the index, and sends updated mappings for expert review. The knowledge remains in a layer that can be updated and inspected without retraining the model.

When Fine-Tuning Is Worth the Investment

Fine-tuning is worth considering when prompting can’t produce sufficiently reliable behavior for a stable, repeated task. The decision depends on what needs to change, whether the improvement can be demonstrated, and whether it justifies the full training lifecycle.

What Fine-Tuning Changes in a Model

Fine-tuning can improve classification, extraction, terminology, output structure, tool selection, and consistency. It’s less suitable for changing information or claims that must link to an original source.

What Fine-Tuning Changes in a Model

These options aren’t mutually exclusive. A model may undergo domain-adaptive pretraining followed by supervised fine-tuning. For LLMs, supervised fine-tuning commonly uses parameter-efficient methods such as LoRA or QLoRA to reduce the computing and memory requirements of adaptation.

Fine-Tuned Content Generation with Runtime Learner Data

A SciForce language-learning platform used fine-tuning to generate structured lessons, personalize exercises, and automate assessments while keeping outputs aligned with curriculum standards.

User profiles and course metadata were stored in application databases. Quiz results, completed activities, and recurring mistakes informed changes to exercise difficulty and future recommendations.

Educators reviewed the relevance of generated lessons, quizzes, and feedback, while students tested personalized exercises and writing and speaking assessments. The team used their responses and performance data to refine generated content and adaptive behavior before a gradual rollout. Continued updates addressed curriculum changes, educational standards, and new user feedback.

How to Prove the Model Improved

A completed training run isn’t evidence of improvement. Evaluation should pass three checks:

  • Data: Are examples representative, consistently labelled, and separate from the evaluation set?
  • Comparison: Does the adapted model beat both the unchanged model and the existing prompted workflow?
  • Regressions: Did better formatting or common-case performance conceal factual decline or rare-case failures?

A 2025 cardiology study shows why evaluation scope matters. Researchers fine-tuned Llama 3.1-8B on real and synthetic discharge summaries, but three cardiologists reviewed only ten selected outputs. The results support feasibility for a defined documentation task, not reliability across institutions or uncommon cases.

The same tests should run after changes to the base model, adapter, dataset, prompt, or serving configuration.

When Fine-Tuning Pays for Itself

Fine-tuning becomes an economic decision when its recurring value can be compared with its lifecycle cost:

When Fine-Tuning Pays for Itself

Value per response may come from shorter prompts, fewer retries, less human correction, or lower latency. Useful lifetime is how long the trained behavior remains valid before labels, requirements, or the base model change.

A high-volume task may still be a poor candidate if it changes every few weeks. A lower-volume task may justify training when each failure requires expensive expert review.

How to Choose the Architecture

A single application may receive requests for exact records, current documents, generated summaries, and repeated domain-specific outputs. Its architecture determines which component handles each request and what happens when the result fails validation.

How to Choose the Architecture

Start with the type of result and its authoritative source. Then decide whether the model needs trained behavior. Define validation and fallback before comparing model sizes, since those execution paths affect both quality requirements and total cost.

These patterns can be combined within one execution path. For example, retrieval may supply current evidence to an adapted generator. A routed workflow operates at a broader level, selecting among generation, retrieval, structured processing, or a sequence of them according to the request.

Route Different Tasks to Different Components

Routing matters when one interface accepts several types of requests. “How many employees joined this quarter?” calls for a direct lookup. “Why did employee turnover change this quarter?” may require records from several systems followed by analysis and a written explanation.

Sending both requests through the same LLM adds unnecessary cost to the first and gives the second no reliable way to access the required data. A router can classify the request, check the user’s permissions, and send it to structured lookup, retrieval, generation, or a sequence of these components.

Routing Enterprise Data Requests

In SciForce’s enterprise data-processing case, the application combined information from HR platforms, CRMs, financial tools, and operational databases. Direct lookups handled requests such as employee statistics and sales figures. Vector search and client-specific knowledge bases supplied business information for source-dependent requests. The LLM was reserved for summaries, trend analysis, and other tasks that required synthesis.

Query-routing logic selected the execution path. Role-based permissions restricted access to client data, while query filters and response validation checked requests and outputs. Simple requests could return data directly; analytical requests could retrieve the required information before passing it to the model.

Model Size as Separate Deployment Decision

Model size should be selected after the workflow has assigned work to retrieval, adaptation, structured processing, and generation. The remaining question is how much model capacity that work requires.

Model Size as  Separate Deployment Decision

A smaller model may perform poorly because it doesn’t have the information it needs, not because it’s too small. Before switching to a larger model, test both options with the same retrieval system, tools, prompts, and checks. If the smaller model meets the quality target without too many retries or fallbacks, a larger model may increase costs without improving the final result enough to matter.

Compare Cost per Accepted Response

AWS Bedrock assigns two Custom Model Units to Llama 3.1 8B with a 128K context window and eight to the 70B version. This gives the smaller model a clear capacity advantage on that platform.

That advantage depends on how many responses pass validation. Retries, calls to a larger fallback model, and human review increase the cost of each usable result. Test both models on the same production sample and calculate:

Compare Cost per Accepted Response

For a local deployment, replace the AWS capacity figures with the required GPUs, peak-traffic replicas, and ongoing maintenance. The cheaper option is the model that delivers an accepted response at the lower total cost.

Conclusion

Suppose a user gets a wrong answer tomorrow. Could you tell within an hour what happened? The system may have retrieved an old document, ignored the right passage, used the wrong model behavior, or sent a simple lookup through generation. You’d also need to know what should happen next: retry the request, check a database, ask for review, or refuse to answer.

Take the request your application can least afford to get wrong and follow it from input to final response. Check where the information comes from, which component handles it, and what catches a failure before the answer reaches the user.

If that path is hard to trace, SciForce can help you review and redesign it.

RELATED BLOG ARTICLES

View all Articles
How to Build Reliable AgTech AI When Farm Data Is IncompleteHow to Build Reliable AgTech AI When Farm Data Is Incomplete

In agriculture, missing data is part of the job. Farm data comes from different sources, under changing field conditions, and at different points in the growing cycle, so a complete and perfectly synchronized dataset is rare. AgTech models still have to work with whatever information is available. Some gaps barely affect the result, while others remove an important part of the signal. Knowing the difference is what makes the model useful outside a clean development dataset. Agricultural data is

# Agriculture
# AI / ML
# Data Science
Scaling AI InfrastructureScaling AI Infrastructure: Navigating GPU Orchestration and Cloud Costs

Microsoft spent $37.5 billion on infrastructure in the quarter ending December 2025, with roughly two-thirds going mainly to GPUs and CPUs. At this scale, even modest waste is expensive. NVIDIA found that idle workloads consumed about 5.5% of GPU capacity across its research clusters. By combining GPU telemetry with job data and automatically clearing stalled or abandoned workloads, it reduced that waste to about 1%, potentially saving millions. Scale down the arithmetic and the pattern holds: a

# AI / ML
# DevOps
Forward Deployed EngineerWhat Is a Forward Deployed Engineer and How to Become One

Forward deployed engineer was once a niche title used mainly by companies such as Palantir. It now appears across OpenAI, Anthropic, Google Cloud, Scale AI, and other enterprise AI providers. FDE brings software engineering, technical consulting, and project delivery into one role. Its growth also shows what enterprise AI companies now need from engineers: an understanding of customer workflows, the ability to make sound technical decisions, and responsibility for moving a system into production

# Tech
# AI / ML
Multimodal AI in Production: Building Systems That See, Hear, and Reason

In 2025, Waymo's driverless vehicles crossed a threshold: independent, peer-reviewed data, not company demos, backed up the safety claims. A peer-reviewed analysis of 56.7 million rider-only miles found a 92% drop in pedestrian injury crashes compared to human-driver benchmarks, plus a 96% reduction in intersection crashes and an 82% reduction in cyclist and motorcyclist injury crashes. Waymo's own dashboard has since tracked the pedestrian figure at 220.6 million miles — more than triple the st

# AI / ML
# Computer Vision
# Speech Processing