MLOps vs LLMOps: What Changes When the Model Is an LLM

MLOps operates machine-learning systems; LLMOps extends those practices to prompts, context, retrieval, providers, tools, evaluations and runtime behavior around large language models.
Published:
Aleksandar Stajić
Updated: October 8, 2026 at 09:31 PM
MLOps vs LLMOps: What Changes When the Model Is an LLM

MLOps is the engineering discipline for reliably developing, deploying, versioning and operating machine-learning systems; LLMOps extends that discipline to applications built around large language models, where production behavior depends not only on a model artifact but also on prompts, context, retrieval, provider/model versions, tool calls, safety controls and evaluation pipelines. LLMOps does not replace MLOps. It changes the operational unit from “a model plus serving pipeline” toward “an evolving LLM application whose behavior emerges from several independently changing components.”

What MLOps really means

MLOps applies software-engineering and operational discipline to machine-learning systems. The production challenge is broader than training a model: data collection, data validation, experimentation, reproducibility, model evaluation, deployment, infrastructure and monitoring all have to work together.

Google's MLOps architecture guidance frames the discipline around continuous integration, continuous delivery and continuous training. CI validates not only code but also data, schemas and models; CD deploys ML pipelines and prediction services; CT can retrain and redeploy models as data or implementations change.

AWS guidance adds the same operational concerns from another angle: model lineage, model/version traceability, drift monitoring and production-quality monitoring are core parts of keeping ML systems reliable after deployment.

What changes when the model is an LLM

Large language models change the production problem because the application often does not own the complete model-training lifecycle. A team may call a hosted model API, run an open model locally, switch between providers or use several models for different tasks.

The model is therefore only one versioned dependency inside a larger behavioral system. Prompts, retrieval results, context order, tools, model snapshot, temperature/reasoning settings, safety filters and runtime orchestration can all change the output.

This creates a broader operational question: which combination of model, context, data, prompt, tools and runtime produced this behavior? LLMOps exists to make that question answerable and the answer reproducible enough for engineering work.

The simplest example

Suppose an application answers internal policy questions.

In a classical ML framing, you might version a trained classifier, deploy it and monitor prediction quality. In an LLM application, the answer might depend on a hosted model snapshot, a system prompt, an embedding model, a vector index, retrieval filters, a reranker and the final selected context.

Changing any one of those components can change the final answer even though the application endpoint and user question stay identical.

A typical LLMOps release path

1
1. Change one component
Prompt, model, provider, retrieval setting, tool schema or application code changes.
2
2. Run deterministic tests
Validate schemas, permissions, tool contracts, retrieval filters and application behavior.
3
3. Run behavioral evals
Compare representative outputs, retrieval quality and agent/tool trajectories against acceptance criteria.
4
4. Compare cost and latency
Measure token use, model calls, retrieval/tool overhead and response latency.
5
5. Deploy controlled version
Ship the concrete application configuration with model/provider versions recorded.
6
6. Trace production behavior
Capture relevant model, retrieval, tool and runtime spans.
7
7. Evaluate production traces
Sample real executions for quality, grounding, safety and task success.
8
8. Roll back or iterate
Use regression evidence and operational signals to decide the next release.

Where the simple example stops

Some LLM systems still train or fine-tune their own models, so traditional MLOps practices such as training pipelines, model registry and data lineage remain directly relevant.

Other systems use only external foundation-model APIs and never run continuous training. Their main operational workload is application evaluation, model/provider change management, prompt/context versioning, retrieval quality and observability.

There is therefore no single universal “LLMOps pipeline.” The exact lifecycle depends on whether you train, fine-tune, self-host, retrieve external knowledge, run agents or depend on managed model APIs.

MLOps vs LLMOps

What stays the same and what expands

MLOpsLLMOps
Primary operational unit
Model ownership
Typical change
Evaluation
Production monitoring
Continuous training
Versioned artifacts
Rollback target

LLMOps extends MLOps rather than replacing it

The core operational principles do not disappear: source control, CI/CD, reproducibility, lineage, deployment controls, monitoring, rollback and measurable acceptance criteria remain essential.

The extension is that more behavior-defining artifacts now sit outside the model weights. A managed foundation model can change behavior through snapshot upgrades, while application output can change through prompt or retrieval changes without any model retraining.

This is why the useful hierarchy is usually DevOps → MLOps → LLMOps/GenAIOps as increasingly specialized operational concerns, not three mutually exclusive practices.

What has to be versioned in LLMOps?

ArtifactWhy it matters
Application codeDefines orchestration, validation, retries and business behavior
Model family + snapshot/versionDifferent snapshots can produce different behavior
Provider / endpointChanges data flow, latency, limits, pricing and availability
Prompt/instruction codeChanges model behavior even with same model
Generation/reasoning parametersCan alter determinism, latency, depth and cost
Eval datasetDefines what “good enough” is tested against
Scorers / gradersDefine how quality is measured
Embedding modelChanges vector representation and retrieval behavior
Chunking/index configurationChanges what can be retrieved
Reranker / retrieval fusionChanges result ordering
Tool schemasChange what the model can request and how
Permission profileChanges what tool actions may actually execute
Context assembly rulesChange what evidence and state reach the model
Safety/guardrail configurationChanges allowed or blocked behavior

Model snapshots become release dependencies

With hosted LLMs, the team may not control model training, but it still controls which model or snapshot the application calls.

OpenAI's current API guidance explicitly warns that prompting behavior can change between model snapshots and recommends pinning production applications to specific snapshots where consistency matters, then running evals when upgrading.

The operational consequence is straightforward: model upgrades should be treated as application releases, not invisible infrastructure maintenance.

Provider lifecycle becomes part of operations

LLM applications often depend on provider rate limits, deprecation schedules, API semantics, context limits, data-handling rules and pricing.

A provider can deprecate a model while your application code remains unchanged. OpenAI's current deprecation schedule, for example, includes 2026 retirement dates for older model snapshots and platform surfaces.

LLMOps therefore needs provider lifecycle tracking, migration testing and fallback decisions in addition to model-quality monitoring.

Prompts behave like production code

Prompts are executable behavioral configuration. Small changes can alter output quality, tool selection and policy interpretation.

OpenAI's current guidance recommends storing production prompts in application code, reviewing prompt changes through pull requests, using typed inputs and covering changes with tests and evaluation checks.

That makes prompt versioning less like editing marketing copy and more like changing a function whose output is probabilistic and model-dependent.

Context engineering becomes an operational concern

The production model rarely receives only a static prompt. It may receive conversation history, retrieved documents, tool outputs, memory, current application state and policy instructions.

LLMOps must therefore observe context assembly: which evidence was selected, which state version was current, whether truncation occurred and whether important instructions survived compaction.

A model regression and a context regression can look identical at the final answer. Tracing the actual context path is what lets the team separate them.

RAG creates its own operational lifecycle

A RAG system introduces a second production pipeline beside model inference: ingestion, extraction, chunking, metadata, embeddings, indexes, retrieval, reranking and context selection.

The knowledge corpus can change every day even when the model and prompt do not. A stale index or broken metadata filter can therefore degrade answer quality without any model drift.

LLMOps for RAG should track corpus/index version, embedding model, chunking policy, retrieval configuration, source freshness and retrieval metrics separately from generation quality.

Evals replace “looks good to me” with release evidence

Generative outputs are often open-ended, so exact-match tests are insufficient for many tasks. LLMOps adds evaluation datasets and scorers that can measure task success, correctness, safety, groundedness, style or domain-specific acceptance criteria.

MLflow's current GenAI evaluation stack supports versioned evaluation datasets, prompt/model comparisons, custom scorers and evaluation over complete traces.

The strongest practice is evaluation-driven development: define representative cases and acceptance criteria before or alongside changes, then compare releases against the same evidence.

LLM-as-a-judge is useful but not ground truth

LLM judges can scale evaluation for qualities that are expensive to encode as deterministic assertions, such as relevance, tone or groundedness.

However, the judge is another model with its own bias, version and prompt. Judge configuration should therefore be versioned and calibrated against human or deterministic reference cases where consequence matters.

A production eval can mix deterministic checks, reference-based metrics, model judges and human review rather than asking one metric to represent every quality dimension.

Tracing becomes more important than endpoint logs

Traditional API logs can tell you that a request took two seconds and returned HTTP 200. They cannot tell you which retrieved chunks were selected, which tool the agent called or which model span consumed most tokens.

MLflow's current GenAI tracing captures prompts, retrievals, tool calls and application spans, and its production evaluation flow can score intermediate trajectory information rather than only final text.

This is a major LLMOps shift: observability follows the behavioral graph of the application, not only the serving endpoint.

Agents expand LLMOps into runtime operations

An agentic application can perform several model calls, tool invocations and state transitions before producing a result.

Operating agents therefore requires step counts, tool-call traces, permission denials, retries, loop detection, human approvals and verified final state in addition to ordinary model latency and token metrics.

A correct final answer can hide a bad trajectory, so agent evaluation must inspect the path as well as the result.

Tokens, model calls and context become cost variables

Classical ML inference cost is often dominated by serving infrastructure or per-prediction compute. LLM applications can add provider token pricing, repeated agent calls, embedding calls, reranking and tool/runtime overhead.

Cost therefore has to be attributed to task or trace, not only to one endpoint. A workflow that makes eight hidden model calls can be functionally correct but operationally unacceptable.

Latency behaves the same way: model latency, retrieval, reranking and external tools compose into end-to-end user latency.

Caching becomes semantic, not only technical

LLM systems can cache prompts, embeddings, retrieval results or full responses, but the cache key must reflect the semantics that can change the result.

A response cache that ignores model version, tenant, permissions or source freshness can return a technically valid but semantically invalid answer.

LLMOps therefore treats cache invalidation as part of model/context/data versioning rather than only infrastructure optimization.

Safety and permissions become release criteria

Generative systems can produce unbounded text and agents can trigger external actions. Safety testing therefore sits closer to ordinary CI/CD than in many classical predictive ML systems.

Permission checks, prompt-injection tests, tenant-isolation tests and side-effect approvals should be reproducible regression tests where those risks exist.

The model may suggest an operation, but the runtime still has to enforce authorization. LLMOps owns the evidence that those controls continue to work after model, prompt or tool changes.

What CI looks like in LLMOps

CI layerExample checks
CodeUnit tests, type checks, schema validation
PromptsTemplate rendering, required variables, policy text, snapshot review
Models/providersCompatibility, output schema, capability and regression tests
RAGChunking fixtures, filter tests, Recall@k, reranker regression
ToolsInput/output schema tests, permission tests, idempotency tests
AgentsTrajectory fixtures, loop limits, handoff/tool-selection tests
SecurityPrompt injection, unauthorized tools, cross-tenant negative tests
Behavioral evalsTask success, correctness, grounding, safety, domain criteria
OperationalLatency, token/cost budgets, timeout/fallback behavior

What CD looks like in LLMOps

A production release may deploy no new model artifact at all. It may simply ship a new prompt, retrieval configuration, tool set or provider mapping.

The release bundle should therefore identify the complete behavior-defining configuration rather than only the application container image.

Feature flags, staged rollout, shadow evaluation, canary traffic and rollback are useful because LLM behavior can regress in ways that static contract tests do not detect.

Continuous training becomes optional; continuous evaluation becomes central

Traditional MLOps often emphasizes continuous training when new data or drift justifies retraining.

Many LLM applications never train the foundation model. Their equivalent continuous loop is continuous evaluation: collect failures and representative production cases, add them to evaluation datasets, test candidate prompt/model/retrieval changes and redeploy only when evidence improves.

Fine-tuning can reintroduce a training lifecycle, but it should sit inside the same broader evaluation and release process.

What should be monitored in production?

Signal classExamples
System healthErrors, timeouts, endpoint availability
Model/providerModel ID, snapshot, rate limits, provider errors
LatencyEnd-to-end, model, retrieval, tool and reranker spans
CostInput/output tokens, embeddings, tool/API spend
QualitySampled task success, correctness, relevance, groundedness
RAGRetrieval recall proxies, empty retrieval, stale sources, citation coverage
AgentsTool selection, retries, loops, handoffs, approval frequency
SecurityDenied actions, prompt-injection indicators, tenant-boundary failures
User feedbackCorrections, abandonment, escalation, explicit ratings
Change driftProvider/model/config changes relative to approved release

Production traces can become evaluation data

One of the most useful modern LLMOps patterns is to turn sampled production traces into evaluation records.

MLflow currently supports retrieving production traces and scoring not only outputs but intermediate spans such as retrieval or tool-call trajectories.

This closes the loop between observability and development: real failures can become regression cases in the next release rather than disappear inside logs.

Reproducibility becomes conditional rather than exact

Classical ML reproducibility often aims to recreate a model from versioned code, data, environment and training parameters.

Hosted LLM applications cannot always reproduce identical output token-for-token because generation is probabilistic and providers may control infrastructure.

LLMOps therefore aims for behavioral reproducibility: record enough model/provider/version, prompt, context inputs, retrieval state and runtime configuration to reproduce the conditions and validate behavior within expected tolerances.

Lineage expands from model lineage to application lineage

AWS's MLOps guidance treats model lineage as the history of code, data, model and infrastructure artifacts needed for diagnosis and reproducibility.

For LLM applications, lineage should additionally connect prompts, eval datasets, retrieval/index versions, tool schemas, agent/runtime configuration and provider/model snapshots.

The target question becomes: Which exact application configuration produced this trace?

Multi-provider and model routing create operational policy

Once an application can use several providers or local models, routing becomes an operational policy rather than a simple model string.

Routing may depend on capability, latency, cost, privacy, context length, availability, tool support or locality. A fallback can preserve uptime while changing answer quality or data-processing assumptions.

LLMOps should therefore log which route was actually selected and evaluate routes independently rather than treat every compatible endpoint as behaviorally interchangeable.

Original implementation evidence

Aaasaasa AI Client: provider, model and runtime are separate operational objects

Aaasaasa AI Client separates agent/client, provider, model, runtime location and permissions. Its AI Hub supports Ollama, LM Studio/OpenAI-compatible endpoints and other provider protocols rather than treating “the model” as one global setting.

The implementation includes dynamic local model discovery, streaming, thinking output and explicit Ollama warm/load and unload controls. That is operational evidence that local LLM serving introduces resource lifecycle concerns beyond an API model name.

Provider status is queried through provider adapters, and connection types distinguish local, cloud API, account-backed, remote-agent and web-client paths. These are concrete operational dimensions an LLM-aware platform has to surface.

The repository also preserves an important boundary: a local runtime is not automatically local inference. Provider/model/runtime location are versioned or configurable concerns that affect privacy, latency, cost and availability.

Source of Truth Research Engine: LLM application state extends beyond the model

The Source of Truth Research Engine combines lexical search, optional embeddings, source snapshots, SHA-256 identity, claims, provenance and contradiction tracking around local model-assisted research.

This is useful LLMOps evidence because changing the model alone does not define the research system. Retrieval, source acquisition, evidence classification and persistent provenance are independent operational artifacts.

The implementation deliberately treats semantic similarity as discovery rather than evidence, showing why LLMOps observability should distinguish retrieval behavior from claim validity.

Observed implementationLLMOps lesson
Multiple provider protocolsProvider identity is an operational dependency
Dynamic model discoveryAvailable models can change independently of application code
Ollama load/unload controlsLocal models have memory/resource lifecycle
Provider health/status adaptersModel availability needs runtime observability
Separate runtime and inference locationDeployment topology is not one boolean “local/cloud”
Central permissionsModel capability and tool authority must remain separate
Lexical + semantic retrieval pipelineRetrieval configuration is part of application behavior
Source/provenance persistenceOperational state and evidence live outside model weights

Common LLMOps failure modes

Failure modeWhat actually went wrong
Model alias upgraded silentlyBehavior changed without controlled release
Prompt changed without evalsBehavioral regression passed normal unit tests
RAG index staleGeneration model was blamed for retrieval/data failure
Only final answer is loggedRoot cause in retrieval/tool/context trajectory is invisible
Provider fallback is silentDifferent model/data path changes behavior without attribution
Token cost tracked globallyExpensive workflows cannot be localized
Judge model changedEvaluation scores drift without application change
Production traces never become testsKnown failures repeatedly return
Local model stays loaded indefinitelyVRAM/resource pressure becomes operational instability
Permissions encoded only in promptModel behavior is mistaken for authorization
One eval score gates everythingDifferent quality dimensions are collapsed into a misleading number
Model registry exists but prompt/index versions do notApplication lineage remains incomplete

Common misconceptions

MisconceptionCorrection
“LLMOps replaces MLOps.”LLMOps extends MLOps principles to LLM-specific application behavior.
“LLMOps is prompt engineering.”Prompts are one artifact among models, providers, context, retrieval, tools, evals and runtime.
“Hosted APIs remove operations work.”They remove some model-serving/training work but add provider lifecycle, version and dependency management.
“If the API is stable, the app is stable.”Model behavior and provider/model snapshots can change independently of API schema.
“RAG is just data preprocessing.”In production it has its own ingestion, index, retrieval and freshness lifecycle.
“LLM outputs cannot be tested.”They can be evaluated with deterministic, reference, judge and human criteria.
“LLM judges are objective ground truth.”They are model-based evaluators that also require calibration and version control.
“A local model eliminates LLMOps.”Local serving adds model files, VRAM, load/unload, runtime health and upgrade concerns.
“Observability means token counts.”Useful observability follows prompts, retrievals, tools, model spans and outcomes.
“Continuous training is mandatory.”Many LLM apps use continuous evaluation without training the foundation model.

A practical LLMOps design sequence

Operate the complete behavior-producing system

1
1. Define the behavior unit
List every component that can materially change output: model, prompt, retrieval, tools, context and policy.
2
2. Establish application lineage
Version code, model/provider, prompts, eval datasets, retrieval configuration and tool contracts.
3
3. Build representative eval datasets
Use expected success/failure cases from design and production.
4
4. Separate deterministic and behavioral tests
Keep schema/security assertions distinct from semantic output evaluation.
5
5. Trace end-to-end execution
Instrument model, retrieval, reranking, tools and agent/runtime spans.
6
6. Define release gates
Set quality, safety, latency and cost thresholds.
7
7. Pin or explicitly record model versions
Treat model/provider changes as release events.
8
8. Deploy progressively
Use flags, canaries or staged rollout where consequence warrants it.
9
9. Evaluate production traces
Measure real task behavior and identify recurrent failures.
10
10. Feed failures back into eval datasets
Turn incidents and corrections into permanent regression coverage.
11
11. Monitor provider and data lifecycles
Track deprecations, index freshness, source changes and runtime availability.
12
12. Retire obsolete versions cleanly
Remove old prompts/models/indexes/credentials after migration and evidence retention decisions.

LLMOps architecture checklist

QuestionExpected evidence
Which model/provider/version served the request?Traceable model identity
Which prompt/instructions were active?Versioned application code/config
Which context reached the model?Context/retrieval trace
Which corpus/index version was used?Retrieval lineage
Which tools were available and called?Tool schema + trajectory trace
Which permissions applied?Runtime authorization record
How is quality measured?Versioned eval dataset + scorers
How are model upgrades tested?Behavioral regression suite
How is production quality sampled?Trace evaluation/feedback process
Can one failure be reproduced approximately?Model/context/provider/application lineage
Where is cost spent?Per-trace model/tool/retrieval attribution
What triggers rollback?Defined quality/safety/cost/availability threshold
How are provider deprecations handled?Migration/fallback process
How are local models operated?Health, resource, load/unload and version controls

Edge cases and limitations

A simple application that calls one fixed hosted model with no retrieval or tools may need only lightweight LLMOps: versioned prompt code, evals, model pinning, basic tracing and provider monitoring.

A self-hosted fine-tuned model may require nearly the full classical MLOps stack plus LLM-specific application evaluation, making the boundary between MLOps and LLMOps intentionally blurry.

An agent platform can have minimal model-training operations but substantial runtime operations because failures occur in tool selection, state and orchestration.

A RAG-heavy system can be operationally dominated by document ingestion and retrieval quality rather than model serving.

Terminology will continue to evolve. The durable architecture question is not which “Ops” label wins, but which artifacts produce behavior and therefore must be versioned, evaluated, observed and governed.

What would change this answer?

If foundation-model providers standardize perfectly stable model behavior and long-term version support, provider/snapshot management could become less operationally significant.

If applications increasingly own fine-tuning or training, classical MLOps concerns become more central again.

The operational principle would remain: every component that can materially change production behavior belongs in lineage, testing, observability and change control.

Related canonical knowledge

LLMOps sits below AI Governance and Enterprise AI Architecture: governance defines which changes require evidence and approval, while LLMOps provides the operational machinery to version, evaluate, deploy and observe those changes.

Context Engineering and RAG are operational subdomains inside many LLM applications because context and retrieval can change behavior independently of the model.

Agentic AI extends LLMOps further into trajectory, permissions and tool-runtime operations.

Frequently asked questions

MLOps vs LLMOps FAQ

What is the difference between MLOps and LLMOps?

MLOps operates machine-learning systems across data, training, deployment and monitoring. LLMOps extends those practices to LLM applications where prompts, context, retrieval, providers, tools and evaluations also materially affect behavior.

Does LLMOps replace MLOps?

No. LLMOps reuses MLOps disciplines such as CI/CD, lineage, evaluation, deployment and monitoring and adds LLM-specific operational concerns.

Do LLM applications need continuous training?

Not necessarily. Many use external foundation models and instead rely on continuous evaluation of prompts, models, retrieval and application behavior. Fine-tuned or self-trained systems can still require training pipelines.

Why are evals so important in LLMOps?

Generative outputs are open-ended and model behavior can change across prompts, snapshots and context. Evals provide repeatable evidence that a release still meets defined quality and safety criteria.

What should be versioned in LLMOps?

At minimum: application code, model/provider/version, prompts, eval datasets/scorers, retrieval configuration/indexes, tool schemas, context rules and relevant safety/permission configuration.

Is prompt versioning enough?

No. The same prompt can behave differently with another model, retrieval set, context order, tool surface or provider.

What is GenAIOps?

GenAIOps is another industry term for operating generative-AI applications. Some vendors use it interchangeably or as a broader label than LLMOps.

How do you monitor an LLM application?

Monitor end-to-end traces including model calls, prompts/context, retrieval, tools, latency, token/cost, quality samples, safety and final task outcomes.

Can local LLMs use LLMOps practices?

Yes. Local models add their own operational concerns such as model files, hardware/VRAM, load/unload, runtime health, quantization and upgrade management.

Glossary

Key MLOps and LLMOps terms

MLOps
Engineering practices for building, deploying, monitoring and maintaining machine-learning systems and their data/model lifecycle.
LLMOps
Operational practices for production applications whose behavior materially depends on large language models and surrounding prompts, context, retrieval, tools and runtime.
GenAIOps
Operational discipline for generative-AI applications; often used as a broader or alternate label for LLMOps.
Continuous training
Automated or repeated retraining and serving of ML models as data or implementations change.
Continuous evaluation
Repeated evaluation of candidate and production AI behavior against versioned datasets and criteria.
Model snapshot
A concrete version of a hosted or packaged model whose behavior can be tested and referenced.
Application lineage
Traceable relationship among code, model/provider, prompts, data/retrieval, tools, runtime and release configuration.
Trace
Structured record of one application execution containing spans such as model calls, retrievals and tool operations.
Eval dataset
Versioned set of representative inputs, expectations and optionally traces/outputs used to measure behavior.
LLM judge
A language model used as an evaluator for qualitative or semantic criteria; it is itself a versioned evaluation dependency.
Behavioral regression
A degradation in application output or trajectory despite interfaces and code continuing to execute successfully.
Provider routing
Policy for selecting among available model providers/endpoints according to capability, cost, latency, privacy or availability.

Conclusion

MLOps and LLMOps share the same engineering objective: make AI systems reproducible enough, testable enough and observable enough to operate reliably in production.

The difference is the shape of the system. Classical MLOps often centers on training and serving model artifacts; LLMOps must operate a behavioral stack in which model snapshots, prompts, context, retrieval, tools, permissions and providers can change independently.

The shortest useful rule is: version, evaluate and observe everything that can materially change the LLM application's behavior — not only the model.

Primary sources and current documentation

The sources below ground the MLOps baseline and the current operational patterns for LLM and agent applications. Project sections are original implementation evidence and are intentionally narrower than claims about a complete LLMOps platform.

Google Cloud — MLOps: Continuous delivery and automation pipelines

Reference architecture describing CI, CD, continuous training, model registry, metadata, serving and monitoring for ML systems.

AWS Machine Learning Lens — Model lineage

Current guidance for tracking code, data, models, environments and infrastructure across ML releases.

AWS Machine Learning Lens — Model observability and tracking

Current guidance for production model monitoring, drift, endpoint health and lineage.

Microsoft Azure — GenAIOps / LLMOps lifecycle

Official guidance describing GenAIOps, sometimes called LLMOps, across initialization, experimentation, evaluation/refinement and deployment.

MLflow — Agents and LLM applications

Current GenAI operations documentation covering tracing, evaluation, prompts and production observability for LLM applications and agents.

MLflow — Evaluating production traces

Current guidance for evaluating complete LLM/agent traces, including retrieval and tool-call trajectories.

MLflow — Evaluating prompts

Current prompt/model evaluation workflow using versioned prompts, datasets, scorers and traces.

OpenAI API — Versioning and model snapshots

Current API guidance recommending pinned model versions and evals because prompting behavior can change between snapshots.

OpenAI — Prompting

Current guidance to treat production prompts as application code, version them through source control and cover changes with tests and evaluation checks.

OpenAI — Deprecations

Current provider lifecycle evidence showing model and platform-surface retirement as an operational dependency.

OpenAI — Moving evaluation workflows to Promptfoo

Current 2026 migration guidance illustrating why evaluation assets should remain portable as provider tooling changes.

Related Articles

Comprehensive Guide to Evaluation Harness: Mastering LLM Performance Evaluation

Comprehensive Guide to Evaluation Harness: Mastering LLM Performance Evaluation

This guide provides a detailed walkthrough of Evaluation Harness, an essential framework for rigorously assessing large language model (LLM) capabilities in enterprise LLMOps pipelines. Learn setup, best practices, and advanced techniques to ensure reliable model benchmarking and optimization.

New Qwen 3.5-Plus: Open-source AI is getting serious now

New Qwen 3.5-Plus: Open-source AI is getting serious now

Discover the groundbreaking features and benefits of Alibaba's Qwen 3.5-Plus, a revolutionary open-source AI for developers.

What Is Context Engineering? What the Model Receives Before It Answers

What Is Context Engineering? What the Model Receives Before It Answers

Context engineering designs what information an AI model receives before inference, including prompts, retrieval, memory, application state, tool results and conversation history.

Mastering the SEO Workflow: Essential Optimization Strategies for Organic Growth

Mastering the SEO Workflow: Essential Optimization Strategies for Organic Growth

A structured SEO workflow is crucial for sustainable organic growth. Learn the ten foundational strategies, from keyword research and technical optimization to content quality and performance analysis.

ZBT Z8102AX Dual-SIM Failover: What Works, What Is Missing and What Needs Better Firmware

ZBT Z8102AX Dual-SIM Failover: What Works, What Is Missing and What Needs Better Firmware

The ZBT Z8102AX is a dual-SIM 5G OpenWrt router, but dual-SIM hardware alone is not the same as intelligent failover. The router recognizes the SIM and connects successfully, but automatic switching, modem recovery, signal-based decisions and clean failover logic still need deeper testing.

Managed Agent Harness vs Self-Hosted Agent Loop: What You Gain, What You Lose

Managed Agent Harness vs Self-Hosted Agent Loop: What You Gain, What You Lose

“Self-hosted agent” can mean very different architectures. This guide separates the managed harness, self-hosted execution environment, and fully self-operated agent loop—and shows which control boundary teams actually need.

Where Does an LLM Get Its Data? RAG Data Sources in Python

Where Does an LLM Get Its Data? RAG Data Sources in Python

An LLM does not magically know your files, databases or APIs. This practical continuation of the RAG series shows, with simple Python, how external data becomes retrievable evidence: from text files and SQL to full-text search, embeddings, context assembly and the final LLM call.

When Should an AI Stop Trusting Its Own Knowledge? — The Retrieval Trigger

When Should an AI Stop Trusting Its Own Knowledge? — The Retrieval Trigger

An AI model does not need retrieval for every question. The important problem is knowing when its internal knowledge is no longer enough. The Retrieval Trigger is a practical decision boundary that determines when an AI system should stop relying solely on model knowledge and obtain external evidence before answering.

What Is RAG? The Simplest Explanation of How It Works

What Is RAG? The Simplest Explanation of How It Works

RAG sounds complicated, but the idea is simple: before an AI answers, it first looks up useful information from a knowledge source and gives that information to the language model. This guide explains RAG, LLMs, state, memory and tools using one simple mental model.

Ollama Is Not the Product: Building Production-Ready Open-LLM Applications

Ollama Is Not the Product: Building Production-Ready Open-LLM Applications

Running a local model with Ollama is easy. Building a production-ready Open-LLM application is harder: it requires RAG, access control, provider abstraction, evaluation, logging, deployment discipline and a controlled application layer around the model.

git-with-automatic-upload-and-synchronization-to-a-production-server

git-with-automatic-upload-and-synchronization-to-a-production-server

RAG Failed — But Which Layer Actually Failed? A Diagnostic Method

RAG Failed — But Which Layer Actually Failed? A Diagnostic Method

When a RAG answer is wrong, blaming retrieval or the model is too vague. This diagnostic method isolates source coverage, query construction, retrieval, ranking, context assembly, generation, evidence attribution, and freshness—so the actual failure can be reproduced and fixed.