2026-09-28 23:31:11
AI SREs are here. If your telemetry, alerts, and access controls aren’t ready, expect missed incidents, alert fatigue, and AI-driven investigations that go nowhere. This Datadog eBook covers how to prepare your stack for the AI SRE future, including a practical self-assessment checklist and real examples from Datadog’s own implementation.
You’ll learn how to:
Assess whether your telemetry coverage gives an AI SRE enough signal to investigate effectively
Design alerts and monitors that an AI agent can act on, not just surface to humans
Set up the access controls and data quality foundations that make AI-driven investigations trustworthy
For 30 years, every online business has been built with a simple assumption that your customers are human. Humans browse the website. Humans enter credit card numbers. Humans click buy.
AI Agents are starting to change this assumption. If ChatGPT is going to be your customer, the internet may need a new protocol.
That’s the question behind Machine Payments Protocol (MPP). To better understand this shift, I attended a MPP event at Stripe HQ, where I spoke with Emily Sands and Matt Schulman from Stripe, along with Brendan Ryan from Tempo.
In this deep dive, we’ll explore:
The problem with today’s payment flow
What is MPP
How MPP works
Two use cases of agentic payment
Your Next Customer Might Be an Agent
Right from its inception, the internet has run on the basic assumption that it will be used mostly by humans. However, this assumption is no longer valid.
According to a recent report from Cloudflare, automated systems are now generating around 57.5% of HTTP requests to web content. AI tools are evolving from question-and-answer chatbots to autonomous agents that can make comprehensive plans, execute actions based on those plans, and evaluate the outcomes of those actions.
However, the internet’s payment model is still stuck in the human-centric era.
This is where MPP (Machine Payments Protocol) comes into the picture. It is a protocol that removes the friction between the buyer and the seller. This friction exists at the interface, which has been built keeping in mind a user reading a web page to figure things out. MPP creates a similar interface for AI agents so that they can settle payments with a service on the internet, irrespective of the payment method.
In other words, the innovation that MPP brings to the table is not only about making payments. It is more about the disappearance of the human from this process.
Agents can generate code. Getting it right for your system, team conventions, and past decisions is the hard part. You end up wasting time and tokens in the correction loops.
More MCPs, rules, and bigger context windows give agents access to information, but not understanding. The teams pulling ahead have a context layer to give agents exactly what they need for the task at hand.
Join us for a FREE webinar on Oct 7 to see:
Where teams get stuck on the AI maturity curve and why common fixes fall short
How a context layer solves for quality, efficiency, and cost
Live demo: the same coding task with and without a context layer
If you want to maximize the value you get from AI agents, this one is worth your time.
If we take the human out of the payment flow, something has to fill the gap. MPP, or Machine Payments Protocol, is the thing that fills this gap.
MPP was co-authored by Stripe and Tempo, and launched on 18th March, 2026. The core of the MPP specification has been published as the Payment HTTP Authentication Schema and has been submitted to the IETF standards track. This is the same body that also maintains HTTP.
What exactly does MPP do?
Before any money moves between two parties, they need to agree to a few things. For example, what is the cost of a service, what counts as payment, and what the proof of payment should look like.
As we discussed, humans make all these decisions just by looking. A user opens a page, finds the price, recognizes the checkout button, and knows what to do next.
Now, MPP manages each of these decisions. It also gives a specific name to each of them:
The Challenge is what the server asks for (payment) when someone requests a service.
The Credential is what the client sends back as proof.
The Receipt is what the server returns along with the service delivery.
These three objects (Challenge, Credential, and Receipt) form the core of the protocol along with payment intents around charge, session, and subscription.
However, every website on the internet has a slightly different way of presenting the payment details to the users. In the case of humans, it rarely matters, because people are good at figuring things out. Software cannot do that. It needs the terms in the same place and format every time so that it can parse the details efficiently.
Therefore, MPP places these details in HTTP headers.
The server sends its terms back in a WWW-Authenticate: Payment header. The client returns the proof in an Authorization: Payment header. The server finally confirms with a Payment-Receipt header. The Challenge contains details such as an ID, the amount, the currency, who to pay, which payment method, and the validity of the offer.
The payment methods are governed separately. Each rail (card networks, blockchain, payment processor) can write and maintain its own specification for how it fits the core. This makes the protocol open enough so that it isn’t controlled by individual companies. The protocol itself is free, with no licensing fees for implementing it.
The ultimate goal of MPP is to become the language of internet payments. A service that understands MPP can sell to any agent. An agent that understands MPP can buy from any service. Neither one has to have heard of the other beforehand.
When an agent requests your API or service, your service returns an HTTP 402 response with payment details. The agent authorizes the payment, retries the request, and gets access to the paid resource along with a receipt. The diagram below shows the overall process flow:
Here’s how the process works:
The agent asks for access to a service or an API.
The server provides a price when it receives an access request. It answers 402 Payment Required and places the terms in a WWW-Authenticate: Payment header. This header carries a Challenge, which contains a unique ID, the payment method, the intent, and an expiry. It also has an encoded payment request that holds the amount, the currency, and the recipient address.
The agent checks the details. There is no account creation, form filling, or buying an API key. The agent reads the terms provided by the server. It checks the various details such as the amount, the recipient, the currency, and the validity window from the signed payload.
Then it authorizes the payment based on a limit that was set earlier as part of the agent’s configuration. These limits are part of the signing key rather than in the agent’s reasoning. This is a delegated key with a spending cap per period, an expiry, a list of permitted recipients, and a scope. Often, there is one key per deployment. Each can be revoked individually so that a runaway agent cannot spend beyond the limit.
In the next step, the agent provides the proof of payment. It repeats the identical request for access using an Authorization: Payment header with a Credential. This includes the Challenge ID with a payment payload, which might be a signed transaction, a paid invoice, or a card token. Credentials are bearer instruments that authorize the spending of real money. Servers and intermediaries don’t log them or echo them into error messages for security purposes.
In the last step, the server provides the access. It verifies the proof against the rail and returns the search results. It also attaches a Payment-Receipt header confirming the settlement.
The overarching rule behind this entire exchange is that the server must not perform side effects (such as database writes or calls to other services) for a request that has not been paid for. The unpaid request that triggers the 402 changes nothing except recording the Challenge. Proofs are single-use, so a Credential sent a second time is rejected.
If verification fails for some reason, the server does not return 401. It returns another 402 with a fresh Challenge along with a structured problem document that provides details about the failure, such as payment-insufficient, payment-expired, verification-failed, and invalid-challenge. The 402 status code means the payment barrier is still present. But the AI agents get to learn about the failure and make corrections.
In the case of Stripe, this setup fits quite well. A successful payment creates a PaymentIntent. The money settles into the existing balance of the business in the default currency. The same tooling works to support other features such as tax, fraud, reporting, and refunds. No customer records are created. No need to deal with user accounts. There is only a receipt on the client side and the payment on the server side. The next time an agent needs something, it has to start from zero.
Let us now look at some use cases of MPP.
As we saw in the previous section, every time an agent needs something, it starts from the beginning. While starting from zero every time sounds great, it also has a cost.
Let’s say the cost for a single web search request is just a cent. However, moving a cent across the banking system often costs more than a cent.
Cards charge a flat fee per transaction. Blockchains charge a network fee per transaction and also take time to confirm. These charges don’t shrink even when the payment amount gets smaller. In other words, below a specific threshold, the cost of settling a payment can turn out to be higher than the payment amount itself.
This is the reason why the internet has relied on subscriptions, credit packs, and monthly plans. Services get bundled up until the value goes above the threshold value. However, agents don’t buy in bundles. They might buy constantly. They could make a search request followed by a lookup. Then, they might ask for another page. This could happen thousands of times inside a single task. If we settle each separately, the processing fee makes things economically unviable. If we wait for each one to confirm, the entire time goes into waiting.
To get around this, MPP does not settle every single payment. It uses the concept of sessions.
In this approach, the agent can open a session by putting money aside at the beginning. This could be a deposit into a reserve. The point is that the deposit is committed, but not yet claimed by the seller. Then, every request is paid with a signed IOU. IOU stands for “I Owe You” and is an informal written record or digital acknowledgment that one party owes a certain amount of money to another party. The agent can sign a message saying another tenth of a cent is owed. It then sends it along with the request. The server checks the signature and serves. There is no immediate settlement.
See the diagram below that shows the session-based flow:
These IOUs accumulate as requests pile up. When the session finally ends, the server claims the total amount in a single real transaction. This way, a single processing fee is divided across thousands of requests. The per-request cost falls to nearly nothing. The delay is just the few milliseconds it takes to perform a simple signature check. In other words, sessions make it possible for agents to pay pennies in an economically feasible manner.
Making micropayments becomes extremely easy with MPP. This opens up some new possibilities.
For example, a reader can use MPP to purchase the latest article from publications like ByteByteGo. Michael Blau demonstrated this using Drip. In this setup, the agents pay per use from an attached Tempo wallet without the need to buy subscriptions. Writers get paid behind the scenes even if the amount is just a single cent. Moreover, the writer can receive the money within a few milliseconds.
As MPP matures, more such use cases are definitely going to appear. Some of the exciting possibilities are as follows:
The internet gets a native payment layer. For years, the web was funded by advertising. However, machines that can pay per request make paid access a strong alternative to the typical advertising-driven internet economy.
Agentic workflows can now transact with the world. This opens new possibilities where agents can buy services like compute, data, and so on. AI becomes more of an economic participant.
It opens a new supply side where we now have a reason to build things specifically for machines being the consumers.
Payments become interoperable at the protocol layer and not at the vendor layer.
Despite the many advantages of MPP, there are also some aspects of MPP that might impact the business in unexpected ways. As we discussed, if we take away the old signup flow, human involvement disappears. But there are also other things that vanish along with it.
Traditionally, the signup process helped the seller find out who was buying the service. The whole signup process created a record containing names, company details, and email addresses. It also helped maintain a history of what a particular customer has done before. Other important requirements, such as abuse control, depend on signups. When someone misuses a service, you could disable their account. Managing disputes and refunds also rely on signups for getting information about the user.
Once we remove the account, all of these things have to change.
With a protocol like MPP, the seller only gets a public key. The payment simply proves that whoever sent it controls the key. It doesn’t tell about which company might be behind the request, which end user the agent might be working for, and whether the same buyer has bought a service earlier with a different key. In other words, payment is no longer a form of identification.
The spending limits can help a little, but only with a specific type of problem. A capped key stops an agent from spending more than it should. It doesn’t stop an agent from spending correctly on the wrong things. The payment is valid, signed, and final. Even if the seller wants to offer a better service, there is nobody to contact after the transaction is done.
The commercial side of business will also see changes:
The signup funnel is no longer present because a person won’t visit the landing page and register for the service.
When a user availed the free tier, it gave an indication about a user who could be potentially converted into a paid customer down the line. That sort of data is difficult to obtain.
The sales call provided details about someone the business could call.
This is why identity and disputes are being built as separate layers rather than being attached to the MPP payment flow.
For identity, there is a way for an agent to sign its requests so that a server can determine which automated client it’s dealing with. These use specifications backed by Visa and Cloudflare. The same key can be recognized across a complete workflow without paying again. This way, the seller can decide whether to trust a particular agent operator.
MPP, however, has not defined a refund flow at all. In a specific session, unclaimed money returns by itself. But in a one-off charge, refunds mean the seller has to send funds back to the key that was used to make the payment. A lot depends on the card network or blockchain provider on whether the refund works as expected.
MPP makes it quite clear what it can control. For example, TLS is mandatory. Challenges can expire. Proofs only work once. Servers should never log a Credential or put one in an error message.
There are also a few guardrails that are part of the MPP specification and should be followed:
The first guardrail is to deal with a runaway agent that spends much more than intended. The solution is to have a delegated signing key with a spending cap per period, a fixed expiry, and a list of permitted recipients. Lastly, there should be a separate key for each deployment so that a particular key can be revoked.
The second guardrail is for a dishonest seller who might quote different prices in the description and the signed payload. The MPP specification tells the clients to verify the amount, recipient, current, and expiry from the signed payload and ignore the friendly text description.
The third guardrail is around discoverability. Directories that list available services and their costs are needed for agents to find sellers. However, according to the MPP specification, these directories and lists should only be treated as advisory when it comes to judging cost and quality.
MPP has been in production since March 2026. The volume so far is relatively small. As of August 2026, about 30,000 transactions are MPP transactions. However, the pattern might be similar to the early days of the Apple App Store, where the first year’s revenue said very little about what the platform would eventually become. What matters more here is the new type of customer that just arrived.
References:
2026-09-26 22:31:13
Agents can generate code. Getting it right for your system, team conventions, and past decisions is the hard part. You end up wasting time and tokens in the correction loops.
More MCPs, rules, and bigger context windows give agents access to information, but not understanding. The teams pulling ahead have a context layer to give agents exactly what they need for the task at hand.
Join us for a FREE webinar on Oct 7 to see:
Where teams get stuck on the AI maturity curve and why common fixes fall short
How a context layer solves for quality, efficiency, and cost
Live demo: the same coding task with and without a context layer
If you want to maximize the value you get from AI agents, this one is worth your time.
This week’s system design refresher:
Top 9 Places to Use Jev
LLM, RAG, AI Agent & Agentic AI
Ex-YouTube Engineer Rebuilds YouTube in 45 Mins (Youtube)
MCP vs Function calling
If Claude Code is a burger...
Jev is TypeSafe AI’s first System One Model. It is 100x faster and cheaper than frontier LLMs.
That opens up a lot of use cases people usually skip because the big models are too slow or too expensive for them.
Here are the top 9 places to use Jev instead of an LLM:
Model routing. Jev takes the prompt as input and routes it to a proper LLM.
Guardrails. Detect security risks in the prompt first. Then pass it to the LLM.
Gating tool-calls. Classify agent tool calls into one of their permission categories.
Triage inbox. Classify lots of emails into categories (e.g., spam, urgent, archive).
Reranking. Score passages against a query (prompt) so we can rank them by relevance.
LLM evals. Evaluate an LLM’s output and return a score within a range.
Bulk labeling. Label lots of rows from a huge table fast and cheap (map-reduce style).
Real-time decisions. Use cases where a fast decision is needed in a loop (e.g., trading).
Confidence gate. Classify anything based on confidence and act properly.
The bottomline is to use the LLM for generations and use Jev on the decisions around it.
What are other places to use Jev instead of an LLM?
LLM: An LLM takes a user prompt and generates a response from its learned parameters. It does so by predicting tokens one at a time.
RAG: In RAG, the user query goes to a retriever before the LLM. The retriever fetches relevant information from the indexed knowledge base and passes it to the LLM along with the query. The LLM generates a grounded response, but it does not guarantee correctness.
AI agent: It has a goal and keeps track of the current task state. It can plan about what to do next, call tools, get a result back, observe it, and repeat until the goal is met. This feedback loop allows the agent to observe the results and adjust its actions accordingly.
Agentic AI: Putting one or more AI agents to meet a shared objective is what makes an AI system agentic. It is achieved through planning, tool use, and feedback. You add an orchestration layer, which coordinates work across multiple agents and workflows. Agents read and write to shared task states and can access tools, data, and the environment.
MCP and function calling have a lot in common. This confuses a lot of engineers, so here is a side-by-side comparison.
Both MCP and function calling are mechanisms that allow an LLM to have access to tools. The agent runtime sends the prompt to the LLM, the LLM decides which tool to use and emits a tool call request.
The agent runtime then handles tool call execution and returns the results back to the LLM to continue. The LLM finally produces a final output shown to the user.
What is different between the two is where the function is implemented, and how the agent handles the execution. In local function calling, the functions are implemented locally on the user's machine. The agent runtime executes them and receives the results.
In MCP, the functions can be outside of the user's machine, on some remote server. The agent runtime follows the MCP protocol and calls the corresponding server, the execution happens remotely, and the results are sent back to the agent runtime.
This allows your agent to connect to hundreds of thousands of tools that are implemented publicly and hosted remotely.
Over to you: What are the best MCP resources out there?
Before each model call, Claude Code assembles a context window from 9 distinct sources.
Think of it as a burger, each layer adds something different.
System Prompt: Defines Claude's role, behavior, and tone. This sets the foundation.
Environment Info: Git status, branch info, and current date. Pulled in via getSystemContext()
CLAUDE.md: A four-level instruction hierarchy: managed → user → project → local. Plain-text Markdown, so users can read, edit, and version-control everything the model sees.
Auto Memory: Contextually relevant memory entries prefetched asynchronously. An LLM scans memory-file headers and surfaces up to 5 relevant files on demand.
Path-scoped Rules: Conditional rules that load lazily when the agent reads files
Tool Metadata: Skill descriptions, MCP tool names, and deferred tool definitions.
Conversation History: Carried forward across iterations.
Tool Results: File reads, command outputs, and subagent summaries.
Compact Summaries: When history grows too long, older segments are replaced by model-generated summaries.
Rebuild YouTube with AI, taught by a former YouTube engineer, kicks off on Saturday, September 26. Enrollment closes in 24 hours.
Scope a realistic MVP. Decide which YouTube features to build and break the work into manageable tasks.
Work with agents. Plan changes, review generated code, and recover from bad diffs or sessions that go off track.
Build across the stack. Turn mockups into React pages and build a backend with Postgres, authentication, and video uploads.
Add semantic search and related videos. Use multimodal embeddings and understand how this approach differs from production recommendation systems.
Test and deploy your app. Check features with Playwright, deploy to Vercel, and track watch time in an admin dashboard.
2026-09-25 23:02:53
Rebuild YouTube with AI, taught by a former YouTube engineer, kicks off on Saturday, September 26. Enrollment closes in 24 hours.
Scope a realistic MVP. Decide which YouTube features to build and break the work into manageable tasks.
Work with agents. Plan changes, review generated code, and recover from bad diffs or sessions that go off track.
Build across the stack. Turn mockups into React pages and build a backend with Postgres, authentication, and video uploads.
Add semantic search and related videos. Use multimodal embeddings and understand how this approach differs from production recommendation systems.
Test and deploy your app. Check features with Playwright, deploy to Vercel, and track watch time in an admin dashboard.
This course is part of ByteByteGo Live, giving you 12 full months of access to all current and upcoming courses.
AI Evals in Practice for Engineers and PMs: Six weeks with Manjeet Singh, Senior Director at Salesforce.
AI Cost Optimization: Four weeks with Jeremy Hintz, Engineering Lead at Meta.
AI Engineering Fundamentals: Four weeks with Ali Aminian (Google, bestselling author).
Build with Claude Code: Two day intensive with John Kim, Senior Staff Engineer at Meta.
Build Production Grade AI Systems: Six weeks with Tanya Roosta, Director at AMD, with a PhD from UC Berkeley.
Trust optimized AI Development: Two day intensive with Kent Beck, creator of TDD.
2026-09-24 23:31:06
A web application can appear quite simple from the outside. A user submits a form, the application saves the information within the form, and a screen displays it later when the user wants to see the information. When the information is no longer needed, the information can be deleted from the system.
However, inside a growing system, the same information or data point can exist in several places. For example, a database holds the original record. A cache might store a copy of that original record for faster access. A search engine index can store some fields of that record to make search more intuitive. An analytics pipeline might produce reports from it. Lastly, backups may preserve earlier versions of that record.
Each of these copies has its own purpose, update schedule, and overall lifetime within the system.
The lifecycle of data shows how information can enter a system, become useful, spread to other components, become old or stale, and eventually be deleted. Thinking about this lifecycle helps connect decisions that might otherwise seem separate, such as database design, performance, reporting, storage costs, recovery, and privacy.
In this article, we are going to look at the entire data lifecycle from creation to deletion and the decisions that need to be taken at each step.
2026-09-23 23:30:40
AI is rewriting the rules for data infrastructure. Streaming keeps critical data moving in real time, while AI introduces new demands for how that data is accessed, governed, and acted on. The Agentic Data Summit, just announced for December 9, brings together engineers, architects, and industry leaders building mission-critical data and AI workloads in production. You’ll hear real lessons from real deployments — governing agent access, migrating off Kafka at scale, and running streaming, SQL, and AI on one unified platform. It’s free, virtual, and built for teams working on what’s next in data and AI.
Modern language models are skilled at many tasks. They can write code, summarize documents, and answer questions across many subjects. However, an application may still need something the model does not provide consistently. Its answers might omit important details, the way it classifies information might confuse similar categories, or its responses might ignore a particular expected writing style.
These things don’t mean that the model is not intelligent or doesn’t have knowledge. Often, it has the underlying capability but needs additional help applying it according to specific expectations.
That is the purpose of customizing a model. One way is to follow clearer instructions and provide better information. But when those also leave gaps between the expectation and reality of the model, further training needs to be imparted. Techniques such as LoRA and QLoRA make that training more practical by reducing the resources needed to adapt an existing model.
In this article, we are going to look at the various strategies to customize and fine-tune a model.
Here are the key learning points in brief:
When instructions and information are not enough
How fine-tuning changes a model
LoRA: Learning a smaller set of changes
QLoRA: Reducing the memory footprint of the model
Tuning the techniques into a training process
Disclaimer: This post is based on publicly shared details from various sources. References at the end. Please comment if you notice any inaccuracies.
The natural starting point in customizing a model is prompting.
A prompt describes the task, identifies constraints, and specifies what a good answer should contain. For example, we can add examples of the type of answer we expect to give the model a clearer demonstration of the expected behavior. This approach is called few-shot prompting.
This simple approach can be quite effective. The task of summarizing a document becomes more useful when it specifies which details matter and how the summary should be organized. Many applications need no customization beyond well-designed instructions and prompts.
However, instructions can’t provide information that is missing. If an answer depends on an internal document or a recently updated policy, the application must provide that material explicitly to the model.
Retrieval-augmented generation (RAG) is the main technique for addressing this requirement. It finds relevant information in an external source and includes it in the model’s input. The model can then use that material when producing its answer.
Prompting and RAG both work through the information supplied during a request. They can guide behavior in a big way, but they don’t ordinarily change the model’s learned parameters.
Sometimes, recurring weaknesses remain. A model might still struggle with specialized document categories or keep producing summaries that highlight the wrong details. While adding more examples as part of every request may help, it also increases the amount of input the application must maintain and process.
This is where fine-tuning offers a different option. In fine-tuning, we can train the model on many examples so that the desired behavior becomes more firmly learned as part of the model’s default behavior.
This is not an absolute division between behavior and knowledge. Fine-tuning can also teach facts. However, retrieval remains useful when information changes frequently, or answers must be traceable to an external source.
As you might be aware, a model’s behavior depends on billions of numerical values called parameters or weights. Many of these are weights that influence how information moves through the model’s calculations.
During pretraining, these values are adjusted using large amounts of training data. A text-generating model learns to predict the next token, which is basically a piece of text such as a word, part of a word, or punctuation. This process helps develop broad language capabilities within the model. Models intended for conversation usually receive additional training to follow instructions.
In contrast, fine-tuning continues from an existing model using a more focused dataset. Since fine-tuning builds on top of capabilities that are already present within the model, it reduces the amount of training required for specialization.
A common approach is supervised fine-tuning (SFT). In SFT, each training example contains an input and the response the model should produce. Examples might pair documents with approved summaries, messages with category labels, or programming requests with suitable code.
During training, the model predicts the desired response tokens. The training process measures how well those predictions match the supplied targets. This measurement is called the loss. Backpropagation identifies how the trainable parameters influence that loss, and an optimizer adjusts them to improve the predictions going forward.
For example, a model may initially answer a classification request with a long explanation. However, training examples that repeatedly pair such requests with short category labels encourage the model to aim for more direct responses. The same principle applies to the details provided in summaries or the conventions followed while generating code.
Across many examples, these adjustments can increase the likelihood of desired behavior based on new inputs. Since the saved changes persist, future requests don’t need to include the entire training dataset in every input they provide.
Supervised fine-tuning does improve things, but it has limitations. Writing explicit examples for every possible scenario the model might encounter is impractical. This is where reinforcement learning from human feedback (RLHF) provides further refinement. The process begins with the model generating multiple responses to various prompts. Human raters then rank these responses based on quality, helpfulness, and safety. These rankings train a separate reward model that learns to predict scores human raters would assign to any response.
SFT and RLHF describe how examples teach the model. On the other hand, LoRA and QLoRA describe how that training is carried out efficiently. The same instruction-response dataset can be used with these different approaches.
The more straightforward type of fine-tuning, known as full fine-tuning, allows every model parameter or weight to change. This provides flexibility, but is also quite expensive. Training must store the weights along with the information used to calculate and apply their updates. Also, intermediate results that were produced while processing examples must also be stored.
This is why a model that fits on a GPU for generating answers might require much more memory for full fine-tuning. Longer examples and larger groups of examples processed together only push the resource requirements higher.
However, as we discussed, the model already has a pretty good grasp on the language aspects. Adapting it to a particular task may not require adjusting every parameter independently. This possibility leads to LoRA.
LoRA stands for Low-Rank Adaptation. It is a form of parameter-efficient fine-tuning, which means it reduces the number of parameters that training needs to update.
LoRA keeps the original model weights frozen. They still perform their calculations, but training doesn’t modify them. Instead, small trainable components are attached to selected calculations inside the model.
These additions are called adapters. They learn adjustments that are combined with the original calculations. This way, the overall output generated by the model changes because the model now uses both its existing capabilities and the learned adjustments.
An adapter operates inside the model. It doesn’t wait for a completed answer and rewrite it afterward. Its adjustments influence the internal processing that eventually produces the answer.
These small components are implemented using two compact tables of numbers, called matrices:
The first creates a compact intermediate representation from the input for a calculation.
The second turns that representation into an adjustment that can be added to the original result.
Training changes the values in these small tables.
The compact representation belongs only to the adapter. The original model still performs its full calculation, so the adapter doesn’t have to recreate all the capabilities the model already possesses. The reason this works is that changes often have shared patterns. If we want a model to write concise technical explanations, we don’t need it to relearn grammar, programming, and sentence construction. Much of that ability already exists within the original base model. Training simply needs to adjust how those abilities are applied.
The term “low rank” refers to representing such an adjustment through a limited set of coordinated patterns. These are learned numerical relationships, rather than explicit rules provided by a developer. Therefore, LoRA restricts how the model can change. This restriction saves resources, but it also creates a trade-off. A small adapter may be sufficient for one task and too limited for another.
A setting called rank controls the adapter’s capacity. Higher rank gives it more room to learn varied adjustments, while increasing its size and training overhead. Values such as 8, 16, 32, and 64 are possible starting points. But a higher rank is not automatically better.
At the beginning of standard LoRA training, the adapters contribute no change. Therefore, the combined model starts with its original behavior. Their adjustments develop as training processes the new examples. LoRA can also target selected attention calculations, which help the model relate different parts of its input, as well as other transformations inside its layers. Targeting more locations provides additional opportunities to adapt.
The savings are related to managing updates for a much smaller collection of parameters. The original model is still necessary and performs substantial computation. LoRA makes training more manageable without turning the underlying model into a tiny model.
LoRA solves much of the expense of updating parameters, but the frozen base model still occupies memory. For a sufficiently large model, storing those weights can remain a major obstacle.
QLoRA, or quantized LoRA, addresses this by combining LoRA with quantization.
Quantization stores numbers using fewer bits and a more limited set of possible values. It preserves an approximation of each original weight while discarding some numerical detail. This reduces storage, but it can also introduce errors into the model’s calculations. The overall model structure and parameter count remain the same. Each weight simply has a more compact representation.
QLoRA commonly stores the frozen base weights in 4-bit form. The adapters remain at higher precision, allowing training to make even finer adjustments to their values. Depending on the implementation, adapter parameters and calculations can use 16-bit or 32-bit formats.
We need to distinguish between storage and computation here. A weight can be stored compactly and reconstructed as a higher-precision value when needed for a calculation. This reconstruction creates an approximation, but doesn’t recover the complete detail discarded during quantization. QLoRA uses this approach to perform the base model’s calculations, combines their results with the adapter adjustments, and trains the adapters using the combined output. The original quantized weights remain frozen.
Although training updates only the adapters, the original calculations still influence what those updates should be. Even if we freeze a part of the model, it is not removed from the learning process. Only the weights are not changed.
Since the adapters learn alongside the quantized model, they adapt to that actual combination. They can help recover task performance affected by compression, although they cannot be assumed to correct every quantization error.
For a fixed memory budget, this can make it possible to customize a larger model than ordinary LoRA would allow. The original QLoRA work has shown great savings, but the practical requirement still depends on the model, length, batch size, and training implementation.
Compressed weights are only part of training memory. The system still needs adapters, their update information, intermediate results, and working space. Lower memory use also does not guarantee proportionally faster training.
LoRA and QLoRA are therefore closely related. LoRA reduces how much must be trained, while QLoRA also reduces the space occupied by the frozen base model.
The first decision is to choose a suitable starting model. The model should already perform reasonably well in the required language and task. This is because a model that struggles with basic instructions or lacks the necessary capabilities will not do wonders even after customization.
Before training, we need to establish a baseline using a carefully developed prompt. This provides a clear picture of the remaining problems and a reference for measuring improvement.
Next, we need to prepare examples that demonstrate the desired behavior. For extraction, include messages with different wording, missing information, and details. For summarization, provide varied documents and consistent examples of how it should be structured.
Correctness is important here because training rewards agreement with the supplied answers. If similar inputs receive contradictory labels, the model can receive conflicting guidance. If approved summaries contain unsupported claims, those claims also become part of the behavior being taught.
Next, we need to prepare separate training, validation, and test sets:
Training examples cause parameter updates.
Validation examples help compare settings and choose a saved version of the model.
Test examples provide a final assessment after those decisions.
Duplicate or closely related examples should not leak across these groups.
The training configuration then controls capacity, update size, and resource use. For example:
Rank and adapter placement determine how much adjustment LoRA can learn. This learning rate controls the size of training updates. Excessively large updates can destabilize learning, while very small ones can produce little progress.
Batch size describes how many examples are processed together. Larger batches generally require more memory. Gradient accumulation allows several smaller batches to contribute to an update. This helps a lot when memory is limited.
Another memory-saving option is gradient checkpointing. It keeps fewer intermediate calculation results and recreates some of them when needed. This trades additional computation for lower memory use and can complement LoRA or QLoRA.
Training duration is measured in epochs, where one epoch means one pass through the training dataset. One to three epochs can be an initial experiment, but repeated exposure isn’t always better. Training longer can make the model overly dependent on the examples it has already seen. This problem is called overfitting. A warning sign for this is when performance on training examples improves alongside worsening performance on validation examples. The useful checkpoint may come before the final training step.
The evaluation process should measure the actual task. For example, classification needs correct labels, extraction needs correct values, and summarization needs faithful coverage of important details. Well-formed output alone doesn’t establish correctness.
Also, we need to check the capabilities the application still depends on. An adapter can improve one behavior while weakening another, even though the original weights remain frozen. The combined model’s behavior has changed, so it must be evaluated as a whole.
LoRA training produces adapter weights that can be saved separately from the base model. These files are generally much smaller than a complete model, but they require the compatible base checkpoint to function.
Keeping adapters separate makes it easier to manage different specializations. Where the serving software supports it, the same base model can be used with different adapters for different tasks.
Another option is to merge the adapter’s learned adjustments into the base weights. This produces a standalone customized model and removes the need for separate adapter calculations. The resulting file contains the full model, so it loses the storage advantage of an adapter-only file.
Merging and quantization support depend on the implementation. We need to evaluate the exact version that will be deployed, since changing its numerical representation can affect results.
The application can still use prompts to specify the current task, RAG to supply relevant information, and validation code to check outputs. Fine-tuning improves learned behavior while these other features continue handling their own responsibilities as before.
Fine-tuning is useful when a model has broad capabilities but needs more reliable specialization. It learns from demonstrations and reflects recurring response patterns more firmly in the model’s behavior.
LoRA makes this less expensive by learning compact adjustments while preserving the original weights. QLoRA reduces the memory requirement further by storing those frozen weights at lower precision.
Neither technique replaces good examples or careful evaluation. The goal is a measurable improvement on new inputs, at a cost the application can support. A clear task, a suitable starting model, and representative training data matter as much as the choice of training method.
References:
2026-09-22 23:32:06
How do you optimize AI when performance matters? And how can you apply AI to make performance better (and easier) than ever before? That’s what 30K engineers will explore at P99 CONF. Here’s a taste of the talks you can expect
Sandboxmaxxing at Lovable: Every Prompt Gets a Sandbox in < 1s
The Autonomous Performance Agent: A Netflix Production Story
Lessons Learned from Building Crazy Fast, Open Source Infrastructure for AI Agents
Give the Agent a Cluster: Effective AI for Performance Engineering at Scale
Managing 500 Billion+ Files for AI Workloads
How to Improve Your Cache Algorithm Using AI
Bonus: Registrants get immediate 30-day access to the complete O’Reilly library, and attendees can enter to win 1 of 500 free swag packs.
If you have used voice assistants before, you have probably experienced unexpected interruptions. You talk with the system, and the moment you pause briefly to think of the right term, it starts talking. Then you interrupt it so you can continue. This is a common frustration, because most voice models can either listen or speak, but not both at the same time.
Newer voice models, like OpenAI’s GPT-Live-1, change that by listening and speaking at the same time. The model constantly decides whether it should stay quiet, interrupt, or start talking. This makes the conversation feel more natural with fewer unintentional interruptions.
Under the hood, these systems combine a new generation of voice model architecture with a serving system optimized for low latency. To understand how it all works end to end, we met with engineers on the GPT Voice team, Zahan Malkani and Justin Uberti (who created WebRTC). We thank both of them for sharing the details with us.
In this article, you’ll learn:
The three generations of voice systems, including cascaded pipelines, turn-based end-to-end models, and full-duplex models
Delegating thinking from talking, the core idea behind GPT-Live
The engineering behind the serving system, including the live and async paths
How evaluation is different in full-duplex voice systems
Engineering lessons for building realtime systems and what is next for voice
Shipping agents to production is the easy part. Keeping them reliable, governable, and improving over time is where most enterprise AI programs stall.
How do top teams do it? They use an Agentic Operating Model (AOM), a step-by-step framework for aligning people, process, and technology so enterprise agents improve as they scale.
In LangChain’s latest guide, you’ll learn:
Why AI agents don’t break like traditional software
The engineering stack that covers the entire agent lifecycle
Shifting from “build and deploy” to “operate and continuously improve”
The input to a voice system is the user’s audio. The output should be the response in audio format played back to the user. While the input and output are always audio, what happens in between depends on how we design the system. Voice systems have gone through three generations of architecture:
Cascaded Design
Turn-based End-to-End
Full-duplex Architecture
Let’s examine each in more detail.
The cascaded design chains three separate models. An automatic speech recognition (ASR) model transcribes the user’s speech into text, an LLM then generates a text reply, and a text-to-speech (TTS) model finally reads the reply for the user.
Each model in the design focuses on the task that it is good at. ASR is good at converting speech to text. It does not have the knowledge that LLMs have. So it only focuses on the conversion. The LLM then relies on its capabilities to understand the user’s query and respond to it accordingly. Once the LLM produces the response text, the TTS model just synthesizes it into speech that sounds natural.
Cascaded design works in practice, but it has two main issues. First, information loss. The LLM only sees a transcript, so vocal information like tone and emotions will be unavailable to the model.
Second, the system is complex and slow. Since the three stages run in series, their latencies add up. The user must wait until all three stages are complete. Also, running three models means building, serving, and scaling three separate systems which is complex in practice.
Due to these limitations, the second generation of voice systems was designed: turn-based speech-to-speech models.
The second generation introduces an end-to-end speech model, a single model trained to consume audio as input and produce audio directly as output.
This design solves a key limitation of cascaded systems: it allows the model to take into account vocal information since it processes the audio directly.
While this is a good improvement, the interaction itself remains turn-based. A small model, called a turn detector, still decides when the user has finished, and only then does the main speech model start its job. The turn detector is not new here. The pause detection step in the cascaded design is the same component. Both generations rely on it, and this turn-based design is the source of unnaturalness.
The detector has a challenging job. If it detects too early, it cuts the user off in the middle of a thought. If it decides too late, users experience awkward delays. Interruptions face the same challenge. When users start talking over the voice system, a separate mechanism has to stop the audio and clear the buffers. If the interruption detection is too sensitive, spurious interruptions can be triggered from background noise. If detection is too conservative, interruptions take too long and feel sluggish, while short interjections like “yes” and “no” can be entirely missed. Justin points to this machinery as the reason earlier voice systems felt unnatural.
These models are also expensive to keep updated. When there is a new, more capable pre-trained LLM, it requires a new full speech-to-speech training run on top of the checkpoint. So voice models always lag behind the newest frontier models.
The third generation of voice models, full-duplex architectures, fixes the turn-taking problem.
Full-duplex models are designed so they can both listen and talk simultaneously. The model continuously produces audio tokens. When it should stay silent, it simply produces silent tokens. Similarly, it continuously processes the input audio tokens. When the user is silent, those are just silent tokens. This design removes the turn detector entirely.
To better understand, let’s use an open full-duplex model, Moshi [x], as a reference. Moshi converts audio into discrete tokens, similar to how text is converted to tokens for LLMs. The model processes these tokens as inputs and emits new tokens on a fixed clock, roughly one frame every 80 milliseconds. Silence is simply another token that decodes to silence. So the model just needs to learn the behavior during training and implicitly understand when to listen and when to talk from the training data.
Full-duplex solves the interruption problem and the unnaturalness in conversations, but it creates two new engineering challenges. First, since the model is always running and predicting the next token, serving is costly. Second, the model needs to respond within milliseconds, so it cannot be too large in capacity.
OpenAI launched GPT-Live-1 in July 2026, a family of full-duplex voice models. It is designed around the challenges described above. The voice model stays small and fast so it can respond within milliseconds. It also relies on delegation to perform expensive reasoning while the conversation keeps going.
This section explains how OpenAI managed to build a voice assistant around a full-duplex architecture. We cover the ideas and techniques that make it practical at scale in production.
Normally, a frontier LLM might reason, search the web, and then respond, which can take several seconds. In voice systems, that translates to a few seconds of silence, which is not ideal.
To fix this, OpenAI separates talking from thinking. A voice model handles the conversation with the user. Another capable model performs the necessary reasoning and tool calling for more complex queries. The serving system is also built around this idea. It delegates the request to the capable model when needed, while continuing the conversation with the user.
As the figure below shows, a question like who won last night’s game cannot be directly answered by the speech model using its internal weights, so the voice model delegates it to GPT-5.5. It keeps the conversation going while the search runs, and reads the answer for the user once it is available.
This design meaningfully reduces the trade-off between speed and quality. While the frontier model looks up information, the voice model remains available to continue chatting. Previously, if you wanted smarter answers, you looped in more systems, and the response took longer. If you wanted quick responses, you used a smaller model, and the answers got worse. With one model for talking and another for thinking, the voice system gets both.
Another benefit of this design is modularity. When a new frontier model is available, less engineering work is needed to switch the voice assistant to it.
Serving two models is tricky. Audio frames should be sent every few milliseconds, but a delegation can take seconds. So we cannot have a shared process for both. GPT-Live splits the traffic. The live path carries audio only. It moves audio between the client and the voice model as fast as possible. The async path handles anything other than the audio, which can take longer. This way, a slow tool call may delay the async path but not the live path.
The main requirement of the live path is that audio must move between the user and the model on a fixed clock. The model processes frames as input and produces them as output in real time. So the user hears a delay when it occurs. As Zahan explains, whenever a bottleneck makes the system fall behind, it starts producing audible artifacts.
To meet this requirement, lots of optimizations are needed. For example, the connection must open quickly, the model must keep up with the incoming stream, and frames must be delivered on time. Here are a few engineering techniques OpenAI adopted to keep the live path fast:
Starting a session in one round trip
Having cheap continuous inference
Handing off live conversations between model instances
1. Starting a session in one round trip
When the user taps the voice button, the client must establish a connection before any audio can flow. This requires a few network trips, depending on the underlying protocol. GPT-Live uses WebRTC, which is used in most video calling apps. But a standard WebRTC session takes six steps to establish before a single audio frame can be sent. On a mobile network where one round trip takes 60 milliseconds, setup alone can cost over a third of a second.
To improve this, OpenAI built WARP (WebRTC Abridged Roundtrip Protocol). It relies on the idea that instead of having the steps one after another, we have them happen at the same time. This shrinks the number of network trips to only one.
2. Having cheap continuous inference
To understand what makes full-duplex systems expensive to serve, let’s compare them with a chatbot. In chatbots, a request arrives, and then the model produces tokens. When the response is complete, the model sits idle. A full-duplex model, on the other hand, has no idle time. Audio streams into the model, and the model produces frames continuously. This happens even when the user is mid-sentence or pauses to think.
Continuously sampling a model, many times per second, for every user, is expensive. OpenAI reduces this cost by keeping each conversation loaded on the model. So instead of re-reading the whole conversation on every request, each session stays connected to the model instance that holds the conversation in GPU memory. When a new audio frame arrives, the model only processes that one frame. In addition, common techniques like batching and speculative decoding can be adopted to keep the system fast.
3. Handing off a live conversation between model instances
Keeping the conversation loaded on one model instance creates a new problem. Model instances may become unavailable. They need to stop and start with demand or receive updates. They may also fail occasionally. In regular systems, this is easy to handle since we can send the request to another available instance. In GPT-Live, the session is tied to the instance that holds its conversation in GPU memory.
OpenAI created a managed handoff mechanism. The system prepares a replacement instance and loads the full conversation beforehand. Once the new instance is ready to be used, the switch happens, so the conversation can continue without any interruptions. This is also useful when other operations need to happen. For example, when the context needs to be compacted, the shortened conversation is prepared on a replacement instance while the original keeps talking. The switch then happens the same way.
The async path handles everything that is not audio, like delegations and tool calling. These jobs, by their nature, are large and can take longer to complete. Still, results should come back fast enough that the two models feel like one system. Back to our example session. The user asks who won last night’s game, and the voice model starts a delegation. The voice model can buy a little time by doing things like acknowledging the question or thinking out loud for a moment. But it cannot do anything when an answer takes too long. So everything in the delegation loop, including prompt processing and routing, needs to happen fast.
To cut latency, we should first understand what is causing it. Most of the time is spent on reading the request. A model that receives a request has to process the entire prompt (prefill) before producing tokens. In a long conversation, that alone can take a noticeable fraction of a second. OpenAI’s fix is to do this reading before it is needed. When a user starts a conversation, the server creates an inference session with the frontier model and sends the conversation so far. By the time the first delegation happens, the model has already everything it needs to produce tokens.
A turn-based speech model can be evaluated one turn at a time by sending a request and scoring the response. But with a full-duplex architecture, there are no turns. There is only one continuous stream of audio. This makes evaluation different.
In full-duplex systems, instead of scoring turns, we can evaluate three parts:
Conversational behavior
The health of the stream
Testing the system on real production traffic
The first question is whether the model behaves well in a conversation. This mostly comes down to timing. The model has to continuously decide if it should stay silent, interrupt, or talk. When the user speaks while the model is talking, it has to decide if this is a real interruption or just background noise. These two decisions are called endpointing and barge-in detection.
To evaluate these behaviors, each decision is scored like a prediction. Did the model detect the end of the turn correctly? Did it recognize a real interruption? The mental model for eval is the same. We collect eval data (natural conversations), have it annotated around the dimensions that we want to measure, run inference on it, and evaluate.
The second dimension of eval is whether the stream itself is healthy. In most serving systems, this is answered with a percentile target like p95. If the p95 latency is good, 19 out of 20 requests feel fast. This works fine as long as unusually slow cases are kept very low. A normal user only sends a request once in a while, so a slow one is quickly forgotten.
But p95 is not reliable in full-duplex architectures. Given that the model runs continuously, a p95 event happens every 20 inferences, which translates to several times per minute. Even p99 events show up regularly. So the system has to be engineered around p999. In practice, this means designing for fast recovery, since some slow frames are guaranteed to happen in every session.
The last layer of evaluation checks whether the system works reliably under real traffic. The technique for finding out safely is called a silent launch. Users keep talking to the old system as usual, while a small percentage of voice sessions are routed to the new system. This kind of silent launch catches problems that are normally hard to find. For example, it can find bottlenecks in unexpected places. In GPT-Live’s silent launch, the team found out that a CPU-side service ran out of capacity before the GPUs did. This is something that is generally hard to learn with ordinary tests.
The broad lesson from OpenAI’s GPT-Live is that building for realtime serving is different and more challenging than traditional serving. For example, in voice systems, users hear the worst frame, so average latency is an insufficient metric to monitor. The tail is more important and deserves more engineering. In addition, capacity planning is different in such systems. When a user is having a conversation, the session is occupied the entire time. So capacity should be measured in concurrent sessions instead of per request.
The other lesson is that complexity should be handled inside the model. For example, turn detection used to be a small, separate component, and it was a main source of unnaturalness. Moving that decision into the model behavior, which has the greatest reasoning horsepower, made the conversation more natural and the system simpler. In general, whatever stays outside the model should be kept small and focused on the realtime work. This ensures the system remains simple and maintainable.
OpenAI believes the next generation is voice driving a computer. In the ChatGPT desktop app, it takes screenshots to see what you are working on, starts long-running tasks, and reports back. Along the way, you can ask for status, answer questions, and redirect it mid-task. Zahan says it feels like science fiction.
Some questions are still open. Zahan says it remains challenging for a voice model to delegate tasks to other models while keeping a live conversation with a user running smoothly. The delegation path might tolerate delay, but live conversation is more sensitive. A general-purpose protocol would mean “part of your brain responds slowly.” Justin says voice AI was like in its GPT-3 days. With GPT-Live-1, he puts it at GPT-4, with many interesting problems still ahead.