Blogpost & Paper Reviews

Notes on engineering blog posts and papers I've read, mostly about how companies build ML and LLM systems in production. For each one I write a summary, what I learned, and the questions I had along with how I answered them. This page syncs daily from the Google Doc where I write them.

Filter by tag

Last updated October 7, 2026

2026 Blogpost original

Voice AI is only as good as what it hears

My notes

What I didn’t understand or am unsure about

Question

Answer

“A support associate would glance at the account, check the spelling, and move on.”

How would the support associate know the account information? Are they cross-referencing based on the phone number?

From the post: The post doesn't say how the associate finds the account.

Inference: The account is usually pulled up before verification starts, most often by caller ID, which matches the calling number to the number on file, or by an identifier the caller gave earlier, such as an account number. The phone number is only a lookup key, not proof of identity, because caller ID can be spoofed. That's why the associate still checks the name.

“To evaluate performance under these conditions, we built an internal benchmark using domain-specific customer service audio.”

What are the main performance metrics for this kind of data? How are they defined and measured?

From the post: Only one metric is named: utterance error rate (UER), how often an utterance has a meaning-changing error.

From web search (μ-Bench blog and README): Sierra's open benchmark also reports word error rate (WER), which is substituted, deleted and inserted words divided by the words in a human reference transcript, and 95th-percentile latency to a complete transcript. The blog's case for UER is that WER counts a dropped "uh" the same as a misheard phone-number digit. An LLM judge decides which errors change the meaning.

Inference: UER is the right primary metric for agents. Its weak point is the judge: I couldn't find how often it agrees with human reviewers.

“And unlike other conversations — where agents can infer intent from a “good enough” transcription — verification is binary: you need an exact match.“

What verification are they talking about here, and why is it binary in this case but not in others?

From the post: Identity verification, such as confirming a caller's name with a bank or insurer.

Inference: It's binary because the transcript is compared with a stored value, and the comparison either matches or doesn't: "Kaitlyn" fails against "Caitlyn." In other turns, the transcript goes to an LLM that can understand the caller's intent despite small errors. Systems can loosen the match, for example by ignoring case or matching on sound, but each loosening makes it easier for the wrong person to pass.

“If context shifts mid-conversation, the challenge gets harder.”

Shouldn't we default to using a multilingual transcription provider then? Any major downsides to that?

From web search (μ-Bench blog): No provider was best in every language, and the most accurate one was among the slowest.

Inference: A multilingual model is a good fallback but a weak default, for two reasons. It usually trails a model tuned for one language on that language. And it has to guess which language it's hearing, which is hard on short inputs like a spelled name.

“Standard transcription pipelines process audio in isolation, without awareness of what the conversation is about or what the caller is likely to say next.”

Similar to the way we give LLMs context at the start of their prompt to inform subsequent generation, are there any architectures/models for voice transcription that can do the same thing?

From web search (Deepgram): Deepgram's Nova-3 accepts up to 100 key terms per request to improve their recognition, without retraining.

Background knowledge: Yes, three kinds. Conventional speech recognition can favor a supplied word list, like Deepgram's. Models with a text decoder, such as Whisper, accept a text prompt treated as the preceding transcript, which is closest to your LLM analogy. Speech LLMs, where an audio encoder feeds a language model, take context exactly like an LLM prompt. The shared risk is that the model outputs a favored word the caller didn't say.

“Instead of picking one provider and managing its weaknesses, we built an ensembler that queries multiple providers in parallel and combines their outputs.“

A major consideration in voice AI is latency. If we're introducing an ensemble layer, what happens to our latency?

Inference: Because providers run in parallel, the added latency is roughly the slowest provider you wait for plus the ensembler's own processing, not the sum of all providers. Two design choices keep it small. A deadline lets the ensembler use whatever has arrived and drop slow providers. A fast path accepts the transcript immediately when providers agree, and runs the slower resolution step only where they differ.

“Cross-reference outputs and identify where providers agree or diverge“

What does this mean concretely?

Background knowledge: The classic method, ROVER (1997), aligns transcripts word by word and votes at each position.

Inference: Sierra likely aligns the transcripts to find the specific words where providers differ, such as a name returned as "Caitlyn," "Kaitlyn" and "Katelyn." It then resolves only those words, using information a vote ignores: each provider's accuracy in that language, confidence scores, and the name on file. Giving an LLM the candidates and the context is a plausible way to do that last step.

“Incorporate signals from earlier turns in the conversation.”

What does this mean concretely?

From the post: The only earlier-turn signal the post mentions is the conversation's language. Its name and address examples come from the customer record, not earlier turns.

Inference: The two most useful signals are probably the agent's last question, which predicts the type of answer (a date, a spelled name, digits), and entities already mentioned, such as an order number or product name that's likely to come up again.

“On our internal benchmarks, we have found that ensembling can cut utterance error rate (how often an utterance has a meaning-changing error) by ~25% on average versus the best single provider, and by up to 37% in languages with more headroom for improving transcription.“

How to weigh this performance improvement versus the presumed latency increase?

Inference (the numbers are my illustration): Compare expected time, because errors cost time too. If ensembling cuts errors from 8% to 6% of utterances, and each error costs a 6-second re-ask, it saves about 120 milliseconds per utterance on average. So it pays off if it adds less than that. Failed verifications and transfers to a human cost far more than a re-ask, so the case is strongest on high-stakes turns like verification and weakest on small talk.

“Rather than asking the transcription model to guess from the full space of possible utterances, we feed it context from the conversation.”

Is the context fed to the transcription model or the agent in a post-processing layer?

From the post: To the transcription side, not the agent. The ensembling diagram also shows context going into the ensembler.

Inference: The context is probably used twice: as bias terms for each provider, and by the ensembler when choosing among their outputs. One risk is that biasing toward the value on file makes a near-miss more likely to be transcribed as an exact match, which weakens identity checks.

“For our financial services agents, context-aware transcription improved input verification rates by over 25%”

What is input verification?

Inference: It most likely means checking that a value the caller says, such as a name, date of birth or account digits, matches the record, with the rate being the share of attempts that pass. Two things are unclear: whether "25%" is a relative increase or percentage points, and whether some of the extra passes are callers who shouldn't have passed. The false-accept rate is the missing number.

“Sierra voice agents have improved resolution rates by up to 1%, which translates to tens of thousands of resolutions a week, and reduced major transcription errors by up to 15%.”

Why is the improvement in transcription errors so much higher than the improvement in resolution rates?

Inference: The 15% is a relative cut in errors, while resolution is measured over all calls, and most calls have no major error. Of the errors that do occur, many don't change the outcome because the agent asks again or understands anyway. And most unresolved calls fail for other reasons, such as policies that require a human. As a rough check, if transcription errors caused failures on F% of calls, a 15% cut would recover about 0.15 × F points, so a full point would need F near 7%. That suggests 1% is a best case.

“the system dynamically reconfigures the transcription pipeline — selecting a different ensemble of providers optimized for that language”

How is the ensemble selected?

From web search (μ-Bench blog): Sierra benchmarks 79 locale variants across 42 languages and more than 13 providers internally, and says this is what lets it pick models per language.

Inference: Most likely, the best ensemble for each locale is chosen offline on accuracy, latency and cost, and stored in a lookup table. When the system detects the caller's language, it switches to that entry.

“No amount of ensembling or context injection will produce a reliable transcription from unintelligible audio. In those cases, the best response is to ask for clarification, just as a human would.”

Is this done using log probs or something else?

Inference: Probably several signals rather than log probs alone. Speech recognition models do return per-word confidence scores, their equivalent of log probs, but those are often poorly calibrated. They're likely combined with disagreement between providers, which the ensemble provides for free, and with checks such as a date that's impossible or a value that doesn't match the record.

How would you improve this system if you were building it today?

Inference: I'd make one change at each stage of the pipeline:

  • Routing. Ensemble only on high-stakes turns such as verification and data capture, and use the fastest accurate single provider elsewhere (#9).
  • Decoding. For structured inputs, restrict output to the expected format, such as digits, dates or letters, and accept the phonetic alphabet ("B as in boy").
  • Verification. Keep "what was said" separate from "does it match," log how much each match relied on biasing, and track the false-accept rate (#10, #11).
  • Learning. Store provider disagreements with human-reviewed corrections, use them to train the ensembler, and alert when the disagreement rate jumps, since that often means a provider has degraded.
  • Evaluation. Report end-to-end task success, for example from τ-voice simulated calls, alongside error rates, because #12 shows transcription gains don't translate directly into resolutions.
My notes

What I didn’t understand or am unsure about

Question

Answer

“Picking a pharmacy. How hard could that be?”

Why is Forus involved in picking the pharmacy at all? Isn’t that something Maria would do herself?

From the post: Forus handles the step after the prior authorization (PA) is approved. Maria's preference is one input, and the agent can “request patient preferences.”

Inference: For a specialty drug the choice isn't free. Her plan often requires a specific specialty pharmacy, and the manufacturer may limit which pharmacies can stock the drug. Most patients don't know these limits. Also, the prescriber's side sends the prescription, and Forus does that work for practices. Maria's preference only decides among the valid pharmacies.

“[W]hich eligible pharmacies are in-network can change without warning. Often, we find out when a routing fails.”

No API connection for checking in-network?

From web search: Drug Channels says there is “no master database” of BIN/PCN/Group codes.

Inference: Contracts between specialty pharmacies and PBMs (the companies that process drug claims for insurers) are private, change on their own schedule, and have no public data feed. The most reliable signal is a rejected claim, which is why a failure is how Forus finds out.

“Insurance routing depends on BIN, PCN, and Group. We may only have two of those three codes.”

What’s the difference between these three and why might you only have two of them?

From web search (Drug Channels): The BIN identifies the insurer or PBM processing the claim. The PCN identifies the benefit package within that payer. The Group identifies the specific plan, such as one employer's plan.

Inference: There are two ways to end up with only two codes:

  1. The code doesn't exist. Some processors don't use a PCN, and some cards have no group number.
  2. The code exists but wasn't captured, for example because the EHR has the wrong card, the photo is blurry, or a form field was skipped.

“Two pharmacies may both be valid choices, so accuracy against a single label can mislead.”

Spectrum of validity or binary?

From the post: Validity is yes-or-no for each pharmacy. The issue is that several pharmacies can each be valid, so the right answer is a set.

Inference: It's binary for correctness: a pharmacy is either in network and able to dispense the drug, or it isn't. It's a spectrum for quality: speed, distance, copay and the patient's preference. Re-route rate only measures the binary part, so a valid but slow pharmacy still counts as a success.

“Unresolved cases went to manual review.”

What does manual review entail?

From the post: Forus's clinical experts do it, using the same tools the agent later got: searching pharmacies, reading documents, checking insurance, contacting the patient and scheduling calls.

Inference: A reviewer likely reads the case documents and works out which pharmacies the plan and manufacturer allow. When unsure, they call the PBM or the pharmacy to confirm network status. They check the patient's preference if it matters, then send the prescription. The phone calls are what make it slow.

“14-case if/else cascade for retail drugs and ~30 hardcoded insurance rules for specialty drugs”

What's the difference between retail drugs and specialty drugs? Why would the rules for them differ?

Background knowledge: Retail drugs are cheap, common and stocked almost everywhere. Specialty drugs are expensive, often biologics that need special handling, and payers or manufacturers often restrict them to specific pharmacies.

From the post: The retail rules compare the pharmacy listed in the EHR with past prescriptions. The specialty rules are insurance rules.

Inference: The two rule sets answer different questions. For retail drugs, almost any in-network pharmacy works, so the question is where the patient likes to go. For specialty drugs, the question is which pharmacies the plan allows, and that depends on the insurance codes.

“Our data was tabular with lots of missing values, which is exactly where gradient-boosted trees shine.”

How do they handle text data? Embedding conversions?

From the post: V2 used no free text; all nine features are short structured fields. Free text (letters, notes, patient texts) only came in with V4. LLM extractors pull a pharmacy name out of the text and fuzzy-match it to Forus's pharmacy database.

Inference: The text fields were most likely treated as categories, which XGBoost handles directly, rather than as embeddings. Extraction fits free text better than embeddings because the useful content is one specific pharmacy, and an embedding would blur it.

“We set per-pharmacy confidence thresholds to target >97% accuracy on automated selections.”

  1. How is the threshold determined and why is it different for each pharmacy?
  2. Reasoning for 97% target?

Inference (1): For each pharmacy, they likely picked the lowest score at which accuracy on held-out data stays above 97%. Thresholds differ because the scores aren't equally reliable. A score of 0.8 for a frequently used pharmacy may be right far more often than 0.8 for a rarely used one.

Inference (2): It's a business tradeoff between the cost of a wrong routing (a patient waiting days) and the cost of manual review, probably benchmarked against how accurate humans are. At 97%, about 3 in 100 automated routings are wrong.

“Instead of a frozen feature snapshot, the model uses a live history of routing outcomes, weighted by recency.”

So the routing outcomes are features?

From the post: Not directly. Matching past routings are summarized into features: the top pharmacy, how much of the evidence agrees on it, how fresh it is, how many routings there are, and how much the rules agree.

Inference: The key change is where the knowledge lives. In V2, the link from insurance codes to pharmacy was built into the model, so it only changed when the model was retrained. In V3, that link lives in the routing history, which updates with every routing. The model only learns how much to trust a given body of evidence, and that changes slowly.

“We search every combination of at least two available features for matching successful routings.”

What does this mean exactly? How are the matching successful routings used?

From the post: They take the fields that are filled in and form every combination of two or more of them. Each combination is a “rule,” for example BIN + PCN. A rule's evidence is the past successful routings that match on those fields. Those routings are weighted by age, summarized into features, and scored by XGBoost.

My arithmetic: Four filled-in fields give 11 rules, which matches the diagram's “10 of 11 matching rules.”

Inference: It's a back-off scheme: when a specific combination has little history, broader combinations still supply evidence.

“The powerset approach substantially outperformed approaches that treated missing features as unknowns.”

By what metric?

From the post: The post doesn't say and gives no numbers.

Inference: It's probably re-route rate, or the share of cases that can be automated at the 97% target. Here is why the powerset approach would win. If “missing” is treated as its own value, a prescription with no Group only matches past cases that also had no Group, which splits the evidence for no real reason. The powerset approach ignores missing fields and uses every routing that shares the fields that are present.

“We keep every matching rule.”

What do they mean by matching rules, and what does it mean to keep them?

From the post: A matching rule is a combination of fields that has past successful routings behind it. Keeping every rule means they don't keep only the most specific one; every rule gets scored, because “when broad and specific rules disagree, that disagreement is useful.”

Inference: Throwing away the broad rules would lose two things. First, a specific rule with 3 routings may be less reliable than a broad rule with 900. Second, disagreement between broad and specific rules signals that something is changing.

“The model uses evidence across the matching rules, including how much they agree.”

How do you measure how much they agree?

From the post: The diagram shows agreement as a count, “10 of 11 matching rules favor RxGate,” meaning the share of rules that pick the same top pharmacy.

Inference: Two variants would probably work better. One weights each rule's vote by how many routings it has. The other compares the most specific rule directly with the broadest one, since the post singles out that split as a sign of change.

“we give recent evidence more weight using a time decay function.”

How are the parameters for the time decay function determined?

From the post: The post doesn't say. Its chart shows a weight near 1.0 for the first week, 0.5 around day 25, and near zero by day 90. It frames the choice as a tradeoff between throwing away useful history and keeping stale routings.

Inference: The parameters were most likely tuned by backtesting: replay past prescriptions in time order and pick the curve that gives the lowest re-route rate. The curve could also be tuned separately for payers whose networks change often.

“For each pharmacy, sum the weights of its historical routing cases. The pharmacy with the highest total becomes that rule’s top pharmacy.“

Why not retain the scores for each pharmacy and use those as features?

From the post: They partly do, in summary form: percent match is the winner's share of the weighted total.

Inference: A separate feature for every pharmacy would tie the model to specific pharmacies. A network change would then require retraining again, which V3 was built to avoid. Your concern still holds: keeping only the winner throws away the runner-up, which matters when the winner has just left the network. One fix is to score each candidate pharmacy within a rule. A cheaper one is to add the gap between first and second place as a feature.

“Time-weighted percent match: How much of the weighted evidence agrees on the winner? A rule at 95% is a much stronger signal than one at 55%.”

What does it mean for a rule to be at 95% versus 55%? What are those percentages measuring?

From the post: The post only defines it: “How much of the weighted evidence agrees on the winner.”

Inference: For one rule, divide the weighted evidence for the top pharmacy by the rule's total weighted evidence. At 95%, nearly all recent routings for that rule went to one pharmacy. At 55%, the top pharmacy barely leads, which suggests two pharmacies are both valid or a switch is underway. Percent match ignores volume, so 95% from 2 routings is weak evidence. That's why routing count is a separate feature.

“One XGBoost model scores each matching rule using its weighted results, evidence freshness, routing count, and the number of fields in the rule, along with agreement across rules. The score estimates how likely that rule’s top pharmacy is to be a valid choice.”

What does it mean for an XGBoost model to score a rule? I don’t really get the statement about how it “estimates how likely that rule’s top pharmacy is to be a valid choice”.

From the post: The highest-scoring rule's pharmacy becomes the prediction, and its score becomes the confidence.

Inference: Think of each rule as a witness that names one pharmacy. XGBoost rates how reliable each witness is, based on its features: percent match, freshness, volume, specificity and agreement. Its output is the probability that sending the prescription to that witness's pharmacy will work. It was probably trained on past prescriptions, with each rule labeled 1 if its pharmacy turned out to be valid and 0 if not.

I’m having trouble reconciling the diagram with the description. Clarify how the XGBoost pipeline works and how the rules fit into it.

From the post: The diagram differs from the text in two ways. It lists routing count as one of each rule's four features and shows agreement separately, because agreement is computed across rules. It also draws a single XGBoost box, but the model scores each rule separately.

Inference: The pipeline runs in this order:

  1. Form every combination of two or more available fields; each combination is a rule.
  2. Pull each rule's past routings and weight them by age.
  3. Compute each rule's features, plus agreement across rules.
  4. Score each rule with XGBoost.
  5. Use the top-scoring rule's pharmacy as the prediction.

“Each extracted name is fuzzy matched against our pharmacy database to select a specific pharmacy.”

And then what? Is that the final decision or another feature somehow?

From the post: By V4, the extracted pharmacies fed into an 11-step cascade with a fixed priority order: “Doctor note beats insurance rule. Insurance rule beats routing history.”

Inference: The extracted pharmacy was most likely a decision at its own step, not a model feature. If a PA letter names a pharmacy, that step fires and the cascade stops there. This is the weakness the post describes, because a year-old note can override fresher evidence.

“At this point, our cascade had 11 steps. Signals often conflicted. The waterfall resolved them using the priority order we set. Doctor note beats insurance rule. Insurance rule beats routing history.”

I thought they were using the XGBoost model at this point, not the cascade?

From the post: They used both. XGBoost never replaced the cascade: “To improve the cascade, we added an XGBoost model.”

Inference: Through V4, the model was one step in the cascade, and “routing history” in the priority order most likely means the model's prediction. Each piece worked on its own, but the code that combined them was still a fixed priority list. The agent replaced that combining step.

“Those tools let it search pharmacies, read documents (including insurance cards via vision), check dispensing restrictions, look up insurance details, request patient preferences, schedule phone calls, store a pharmacy selection, or escalate to an expert. It reads everything, decides which pieces of information to rely on, investigates further if it needs to, and then acts.“

Do we have any latency constraints? This feels like a time-consuming process.

From the post: The post mentions no latency limit. Its time scale is days: a wrong pick delays treatment by days.

Inference: The time budget is hours, not milliseconds, because this runs in the background after the PA is approved. A few minutes of agent work is tiny compared with a re-route. The real constraints are cost per case and human wait time, since asking the patient or calling a pharmacy adds real delay. A sensible design is to act right away when the signals agree and investigate only conflicts.

“Ask Maria for updated information. Escalate if selection remains unclear.“

What does escalation mean here?

From the post: Escalation hands the case to a Forus clinical expert, the same people who did manual review earlier.

Inference: “Unclear” probably means Maria doesn't reply, or her answer still doesn't lead to a valid pharmacy. The expert then calls pharmacies or the PBM and decides. Ideally the expert gets the evidence the agent already gathered, so they don't start from scratch.

“The first version of the agent performed worse than the waterfall in many cases. But when it got things right, its reasoning reflected the kind of cross-signal thinking that our clinical experts exhibit.”

How much does the latter matter given the former? Is there any good way to combine the strengths of the agent with the waterfall?

From the post: They didn't ship based on reasoning quality. The agent ran in shadow mode for weeks, and its actions were unlocked gradually.

Inference: Overall accuracy matters more, because patients bear the cost of errors. The good reasoning mattered as a sign of potential. The waterfall's flaw is built into its design, since a fixed priority order can't spot stale signals. The agent's errors, by contrast, can be fixed with better prompts and tools. Three ways to combine them:

  1. Let the waterfall handle cases where the signals agree, and send conflicts to the agent.
  2. Give the agent the V3 model's prediction as a tool.
  3. When the two disagree, investigate or escalate.

“We couldn't A/B test with live prescriptions flowing through our system.”

Why not?

From the post: The post gives no reason.

Inference: An A/B test would deliberately route real prescriptions through an unproven system, and each wrong routing leaves a patient without medication for days. The agent's actions, such as texting patients and scheduling calls, also can't be undone. Shadow mode avoids both problems and compares the two systems on the same prescriptions. Its limit is that when the agent picks a different pharmacy, you can't see whether that pick would have worked.

“A denial from a previously valid pharmacy or a provider note pointing somewhere new can prompt it to request a patient preference, suggest a phone call to verify network status, or escalate to a clinical expert.”

How do you actually encode these conditions in code? Would a system that prompts further digging whenever a high-confidence prediction is wrong help with that?

From the post: The post shows no code.

Inference: It probably works in two layers. Ordinary code detects events, such as a pharmacy rejection or a new note, and reopens the case. The LLM then chooses a response from its tools, within guardrails such as a list of allowed actions.

Your idea would help, and it builds on the post's own trigger. One failure on a confident route should flag that route for every pending prescription and trigger a single verification call, instead of letting each patient fail in turn. That fills the gap the post admits: time decay “doesn't tell us that a network has changed.”

How would you improve this system if you were building it today?

Inference:

  1. Evidence: treat one failure on a confident route as a warning for every prescription on that route (Q25), and check high-volume routes on a schedule.
  2. Decision: let the V3 model handle easy cases and send only conflicts to the agent (Q23).
  3. Evaluation: re-route rate misses prescriptions that stall at a valid pharmacy, so also track the time from approval to first fill.
2026 Blogpost original

How we evaluate our LLM judge: a perturbation-based approach

The verification LLM needs an evaluation strategy. They had a gold standard dataset which contained correct answers but they also wanted to be able to assess how the pipeline behaved for wrong examples.

My notes

Summary

Context

Insurance plans require prior authorizations (PAs) to decide whether to cover specialty medications for conditions like autoimmune diseases, cancer, and multiple sclerosis, based on detailed questions providers answer about patients’ clinical history.

Forus has a pipeline that generates answers to these questions and then verifies them using an LLM that checks patient clinical records and specialty-specific guidance

Problem

The verification LLM needs an evaluation strategy. They had a gold standard dataset which contained correct answers but they also wanted to be able to assess how the pipeline behaved for wrong examples.

Solution

Perturb correct answers to make them wrong in plausible ways and then check performance on them.

Some key insights/new things I learned

  • Perturbation of gold standard data to generate synthetic negatives
  • Decomposition of statements into independently verifiable claims

Design Details

  • Data
    • “gold standard dataset for our PA evals through a multi-annotator consensus workflow followed by expert clinical review, sampled to be representative of our real data distribution across drugs, payers, and question types”
    • Synthetic data generated from perturbations of gold standard dataset
    • Patient clinical records and specialty-specific guidance
  • Architecture
    • Answer to PA form question generated
    • Answer decomposed to independent claims
    • Judge verifies each claim against supporting evidence
      • Evidence retrieved from the patient record using query expansion and semantic search
    • If any claim not supported, answer flagged for review
  • Evaluation
    • Data:
      • Gold standard dataset of correct answers
      • Synthetic (perturbed) dataset of wrong answers
    • Metric(s):
      • FPR for gold standard, detection rate for perturbed data
    • Results: [Not shared]

What I didn’t understand or am unsure about

Question

Answer

“LLM-as-judge systems have a well-known failure mode: they tend to be overzealous, unnecessarily flagging correct answers.”

Any studies where I can find more details about this besides the source they linked?

From web search: Three more targeted sources:

  • VeriFact (Chung et al., NEJM AI 2025) is the post’s own reference [6] and the closest match: an LLM judge checking clinical text against EHRs with the same three labels. It reports the judge “is biased in assigning Not Supported rather than Not Addressed verdicts when compared against humans.”
  • Jing et al., “On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation” (NAACL Findings 2025), found Llama 3.1 8B “overly strict” but Claude Haiku “extremely lenient.”
  • MiniCheck’s LLM-AggreFact benchmark (Tang et al., 2024) compares fact-checkers that verify claims against grounding documents.

Inference: Which way the bias goes depends on the model and the prompt, so “well-known failure mode” overstates it. Over-flagging is common in strict grounding setups like this one, but it isn’t universal.

“every extraneous flag leads to additional steps”

What additional steps typically?

Inference: The likely sequence is:

  • A human reviewer (Forus operations or clinical staff) re-reads the chart to check the flagged claim.
  • If the chart is ambiguous or incomplete, someone contacts the provider’s office for clarification or documentation. This is usually asynchronous and can take days.
  • The answer is corrected or confirmed, possibly with provider sign-off.
  • The form is re-verified and submitted to the payer.

For a false flag, all of this is wasted, because the original answer was right. The main cost is time spent waiting in the queue rather than reviewer minutes, since the PA can’t be submitted until the flag clears. AMA surveys report that PA delays often lead patients to abandon treatment (web). That’s why the post treats false flags as a clinical cost and not just an operational one.

How are the “generated PA answers” being generated in the first place?

From the post: Not described. The post only calls them “generated PA answers” and says the platform “manages PA submission end-to-end.”

From web search: Forus’s site says providers write the prescription in the EHR and “Forus automatically generates and submits prior authorization forms.” Fierce Healthcare reports that each prescription gets its own AI agent that reasons over the patient’s medical history, insurance and finances, “drawing on clinical models, specialized sub-agents and what Forus has learned from millions of cases.”

Inference: The generator is most likely an LLM agent. It pulls the patient’s record through an EHR integration, retrieves evidence for each payer form question, and writes a free-text answer or picks a multiple-choice option. The judge’s retrieve-then-judge design mirrors this. That raises a question the post doesn’t answer: if the generator and judge share retrieval components, they may also share blind spots, such as a note that neither of them retrieves.

“At Forus, we developed an LLM judge to verify generated PA answers”

This is different from standard LLM judges that are used for evaluation, right? More akin to LLM verifier

From the post: Yes. The judge runs on generated PAs to “verify generated PA answers against patient clinical records and specialty-specific guidance,” and its flags send cases to humans (“every flagged case requires human intervention”). That makes it a production gate, not an offline quality score.

Inference: It differs from a typical evaluation judge in three ways:

  • When it runs: online, on every case before submission, rather than offline on a test set.
  • What it checks against: retrieved patient evidence (groundedness), rather than a reference answer or a quality rubric.
  • What its output does: it triggers a routing decision, so a false flag has a direct operational cost. An evaluation judge’s errors only distort a metric.

“Verifier” is a fair term. In the LLM literature, though, that word often means a model that scores reasoning steps or candidate solutions, for example to pick the best of N samples. “Groundedness checker” or “guardrail” is more precise here. This framing also explains why the post evaluates the judge the way you would evaluate a classifier.

Is verification done question by question or in bulk?

From the post: Verdicts are produced per answer. Each answer is decomposed into claims, the claims are judged independently, and “question-level aggregation” rolls them up into an answer-level result. In the Crohn’s example, though, the judge flags an unperturbed answer because it conflicts with another answer, so the judge must be able to see the other answers on the form.

Inference: The unit of verification is one question’s answer, with the rest of the form available as context. The post doesn’t say whether the calls run one at a time, in parallel, or as one form-level prompt. My guess is separate per-answer calls (which can run in parallel), with the form’s context included during decomposition, since the post says decomposition “has to draw on question context, not just the answer.”

“Our system faced a calibration problem”

If they didn’t have a dataset initially, how did they know this for? Or did they assume it was the case given it’s a problem generally?

From the post: The premise is partly off, because they did have a dataset. They had already built “a rigorous gold standard dataset for our PA evals” of correct answers. Running the judge over correct answers and counting flags gives the false positive rate directly, and that is exactly the clean run. What they lacked was examples of wrong answers on the same patients, which is what perturbation adds. The post also points to industry findings [1] that judges lean conservative.

Inference: The post doesn’t say how they first noticed. The most likely route is that human reviewers in production kept clearing flags on answers that turned out to be fine, or that an early clean-run measurement showed a high flag rate. Note that “calibration” here means the judge flags too aggressively relative to the true error rate. It isn’t probability calibration in the statistical sense, because the judge outputs discrete verdicts, not probabilities.

“Emerging work like MedAgentBench [5] is more targeted, and evaluates LLM agents in simulated EHR environments. Still, no standardized benchmark exists for PA question answering.”

What's the difference between these two use cases?

From web search: MedAgentBench (Stanford, NEJM AI 2025) has 300 physician-written tasks across 100 de-identified patient profiles in a simulated EHR that follows the FHIR data standard. Tasks are either queries (retrieve information) or actions (for example, order a medication), and models are scored on task success rate.

Inference: The two differ in three ways:

  • Output: MedAgentBench asks for one fact or one action with a single correct result. PA question answering asks for answers to payer questions (diagnosis, prior therapies, severity) that must be justified by chart evidence.
  • Data: MedAgentBench draws on structured records (labs, vitals, procedures, diagnoses, orders). PA answers often rest on unstructured progress notes, biopsy reports and outside records.
  • Grading: MedAgentBench checks whether the task succeeded. PA answers need every claim grounded in the record, and they must match payer-specific criteria.

In short, MedAgentBench tests whether an agent can operate an EHR, while PA question answering tests whether it can build a defensible case from one.

“rigorous gold standard dataset for our PA evals through a multi-annotator consensus workflow followed by expert clinical review”

Are the annotators different from the expert clinical reviewers? Why? Cost?

Inference: This is a standard tiered labeling design, and each tier handles a different kind of error:

  • Random error: several annotators label each item, and taking the consensus removes individual slips. Trained non-physician staff, such as nurses or PA specialists, can do this at lower cost and higher volume.
  • Systematic error and hard cases: experts (physicians or pharmacists) review the disagreements and a sample of the agreed labels. This catches mistakes that all the annotators make together.

So yes, cost is a large part of it, because expert time goes only to the cases that need it. Expert adjudication would be needed anyway, since even clinicians often disagree on this task. VeriFact reported that the best agreement between any two clinicians was 88.5% (web).

If an answer is flagged for review, what happens next exactly?

Inference: A plausible flow is:

  • The flagged answer goes to a human review queue, together with the failing claim and the evidence the judge cited.
  • The reviewer checks the chart and either confirms or overturns the flag.
  • If the flag is confirmed, the reviewer corrects or regenerates the answer, and asks the provider for documentation if the chart is missing something.
  • The corrected form is re-verified and submitted.

Claim-level verdicts make review faster, because the reviewer checks one failing claim rather than re-verifying the whole answer. The reviewers’ decisions are probably logged, and those logs are a natural source of the “confirmed errors” the post wants to mine (see Q25).

“each answer is decomposed into atomic claims”

Is this done by an LLM? If so, is it the same kind of LLM as the generation model?

From the post: It isn’t stated outright, but it’s strongly implied. The post notes that one claim (“Humira was the most recent biologic therapy”) “appears nowhere in the answer” and is implied by the question, so “decomposition has to draw on question context.” Rule-based sentence splitting can’t produce a claim like that. The method is also adapted from VeriFact and FActScore.

From web search: Both of those use an LLM for decomposition. VeriFact used Llama 3.1 70B to extract “simple sentences that encapsulate Subject-Object-Predicate relationships.”

Inference: It is almost certainly an LLM. The post doesn’t say which model, or whether it matches the generator. Decomposition is a fairly mechanical task, so a smaller and cheaper model would be a reasonable choice, with the strongest model saved for the judgment step.

Is the judge model the same as the generation model?

From the post: Not stated.

Inference: There are two reasons to prefer a different model, ideally from a different family. There’s no evidence either way about what they did.

  • Self-preference: LLM evaluators tend to rate their own outputs more favorably (Panickssery et al., 2024, “LLM Evaluators Recognize and Favor Their Own Generations”; background knowledge).
  • Correlated blind spots: a model that misreads a chart pattern while generating may misread it the same way while verifying.

One point cuts the other way. Their main problem is over-flagging, not leniency, so self-preference doesn’t seem to be their main failure. Their remark that “using another LLM to judge the judge is circular” shows they are aware of correlated errors in general. The same independence concern applies to the perturbation model (see Q21).

“We expand each claim into multiple semantic search queries”

How is the expansion done? SPLADE?

From the post: The method isn’t named. The example queries are full natural-language paraphrases, such as “symptoms well controlled on Humira” and “stable disease on anti-TNF therapy.” They vary clinical synonyms, brand and generic drug names, and drug class.

Inference: This looks like LLM-generated query rewriting (often called multi-query retrieval), not SPLADE. SPLADE is a learned sparse encoder. It turns one text into a single sparse vector of weighted vocabulary terms, including related terms that weren’t in the text. It doesn’t produce several readable query sentences, so you wouldn’t describe its output as “multiple semantic search queries.” The results of the separate queries are then presumably merged, for example by taking the union and removing duplicates, or by reciprocal rank fusion. The post doesn’t say which. Phrasing several queries around the same claim also makes it more likely that contradicting notes get retrieved, like the April 2026 flare note in their example.

Is their search only semantic or includes lexical as well? Why?

From the post: Only “semantic search queries” are mentioned. Lexical search isn’t.

Inference: We can’t tell, but a hybrid design would make sense, because each method covers a weakness of the other:

  • Semantic (dense) search handles the synonym problem the post emphasizes, such as “eczema” versus “atopic dermatitis.”
  • Lexical search (for example BM25) handles exact tokens that dense embeddings represent poorly. These include ICD codes (K50.90), lab values (CRP 12 mg/L), doses (90 mg every 8 weeks), dates and rare drug names, and PA questions depend heavily on exactly these.

VeriFact, which their pipeline is adapted from, compared dense search, hybrid dense-plus-sparse search, and hybrid search with a cross-encoder reranker (web). Two other points apply. Structured data such as medication lists and labs may be fetched by direct lookup rather than search. And the LLM query expansion partly substitutes for lexical recall, because it writes both brand and generic names into the queries.

Any domain adaptation for semantic expansion?

From the post: Nothing beyond the examples. They show clinical synonyms (“eczema” for “atopic dermatitis,” “failed therapy” for “inadequate response”) and drug-name variants (Humira, adalimumab, anti-TNF therapy).

Inference: Domain knowledge could enter at three points, none of them confirmed:

  • Query generation: an LLM prompted with clinical instructions and examples. This is the cheapest option and the most likely.
  • Vocabulary normalization: mapping terms to medical ontologies, such as RxNorm for drugs and SNOMED CT or ICD-10 for diagnoses, so that “Humira” and “adalimumab” resolve to the same concept.
  • Embeddings: a biomedical retrieval model such as MedCPT, or a general embedding model fine-tuned on their own claim-to-evidence pairs, which their gold standard data could supply.

My guess is that prompting does most of the work, possibly with an ontology layer for drug names.

What granularity/chunking strategy for retrieved evidence?

From the post: Not stated. The diagrams show evidence as short excerpts labeled with source type and date, such as “Progress note · Apr 2026: ‘flare; symptoms recurring’,” “Lab results · Apr 2026” and “Medication list.”

Inference: That suggests sentence- or passage-sized chunks that carry metadata (document type and date). Structured data such as labs and medications is probably stored as records rather than text. Dates matter because many PA claims are about time, such as “most recent biologic” or “sustained remission.” The judge can only see that an April 2026 flare contradicts a January 2026 remission note if each chunk carries its date.

From web search: VeriFact, their reference design, split notes into 128-token chunks along semantic boundaries, then broke those into sentences or atomic propositions. It retrieved the top N facts per claim. It tested N up to 50, and performance was still improving at 50.

“Not Supported: the claim conflicts with the evidence, either through direct contradiction or through a significant omission or fabrication.”

How is the omission part different from Not Addressed?

From the post: Not Addressed means “the evidence is silent on the question.” The judge’s rationale for the adherence example, which appears when you click the verdict, draws the line: silence “isn’t fabrication either, since the claim asserts no specific finding the chart should contain.”

Inference: The difference is whether the record says anything on the point:

  • Omission: the record does contain relevant information, and the answer leaves out the part that changes its meaning. For example, the answer says “methotrexate was tried” when the record says it was stopped for intolerance. That matters for a “tried and failed” question.
  • Not Addressed: the record contains nothing on the point, so there is nothing to omit.

Fabrication sits between them. The claim asserts a specific finding, such as a lab value, that the chart would contain if it were true, and the chart doesn’t contain it. This boundary is hard even for people. In VeriFact, clinicians agreed only about a third of the time on Not Supported labels, and the LLM judge leaned toward Not Supported over Not Addressed (web).

“Worst-verdict aggregation means the one Not Supported claim flags the whole answer. But the Not Addressed claim does not”

So then what happens if there is a Not Addressed claim? Shouldn’t that still go to additional review?

From the post: A Not Addressed claim doesn’t flag the answer, and “that’s the distinction that keeps the judge from over-flagging.” Their reasoning is that records are routinely incomplete, so treating silence as Not Supported “would punish every incomplete chart.” The post doesn’t say what happens to these claims afterward.

Inference: Your concern is valid. The answer goes to the payer containing a claim nobody can document, and payers can deny or audit based on documentation. A middle path is to route these claims by how much they matter:

  • Claims the payer’s criteria depend on, such as prior therapies or severity scores, go to a lighter “documentation gap” queue, or the provider is asked for the documentation.
  • Incidental claims, such as adherence when the question didn’t ask about it, pass but get logged.

There’s also a separate question: why is the generator asserting things the chart doesn’t mention? The share of Not Addressed claims is worth tracking as a signal of generator quality, even if it never triggers a flag.

“the perturbed run measures the detection rate”

Detection rate = true negative rate?

Inference: It depends on which class you call positive:

  • If a wrong answer is the positive class and a flag is a positive prediction, which is how the post frames it, then detection rate is the true positive rate for errors (also called recall or sensitivity).
  • If you call a correct answer the positive class, detection rate becomes the true negative rate (specificity).

Both readings are internally consistent. The post uses the first, which is the standard one for error detection; under that framing, “false positive” means a correct answer that was flagged. Two details about the denominator: accidentally correct perturbations are excluded from it, and only direct detections count, with consistency-based flags tracked separately.

“because detection rate is only meaningful once we know the judge isn't excessively flagging”

Why?

Inference: Detection rate alone rewards flagging more. A judge that flags every answer has a 100% detection rate. A judge that flags answers at random with probability p has a detection rate of p, with no skill at all. So detection rate only shows skill as the gap above the false positive rate. For example, 60% detection with a 40% false positive rate means the judge barely tells wrong answers from right ones, while 60% detection with a 2% false positive rate is a useful judge. The two numbers have to be read as a pair, like recall and precision.

There’s a second reason. With a high false positive rate, some flags on perturbed answers may be for reasons unrelated to the perturbation, so the detection count overstates real catches. In principle, their claim-level rationales could be used to check that each flag points at the perturbed claim. Anchoring on the false positive rate also reflects their priority, since over-flagging is the failure they set out to fix.

“One perturbation, two flags, and the second one is, interestingly, a behavior we hadn't explicitly designed the judge to produce.”

I don’t fully understand the difference between these two flags.

From the post: The interactive demo labels the two flags explicitly:

  • Flag 1, “direct detection,” is on Q1, the answer they changed. The diagnosis now says ulcerative colitis, but the chart documents Crohn’s disease, so the answer contradicts the record.
  • Flag 2, “consistency-based,” is on Q2, an answer they didn’t change. “Stricturing disease with perianal fistulas; transmural inflammation” is still true according to the chart. But those features are hallmarks of Crohn’s, not ulcerative colitis. The demo’s rationale says Q2 “no longer fits the diagnosis it’s meant to support, so the judge flags the mismatch.” Perturbing Q2 instead produces the mirror case: Q1 gets flagged “even though this answer is chart-supported.”

Inference: Flag 1 checks an answer against the record, and flag 2 checks an answer against another answer. Measured against the chart alone, flag 2 is a false positive. It’s a useful one, though, because the submitted form as a whole would no longer make sense. That’s why the post counts the two kinds separately.

“Our approach was to build a perturbation model, seeded with structured context from the patient's record, so its wrong answers stay anchored to the patient's actual medical history. For free-text questions, the model generates a plausible wrong answer.”

So they’re just giving the model relevant context and asking it to generate answers that seem plausible but are wrong? Any configurations to think about beyond the context and the prompt?

From the post: Essentially yes. The model gets structured context from the record. For free-text questions it writes a plausible wrong answer. For multiple-choice questions it picks a different option that is “clinically coherent in context but factually unsupported.” The constraints are definitive assertions, specific and falsifiable claims, no attestations, and about 30% of questions perturbed per form.

Inference: Other settings worth deciding, grouped by stage:

  • Before generation: which questions to perturb (at random, or stratified by question type), and which error type and level of subtlety to aim for. Ideally these come from an error taxonomy (see Q26).
  • During generation: which model to use (ideally from a different family than the judge, to avoid shared blind spots), and how many candidates to sample per question.
  • After generation: validity checks (is the new answer actually contradicted by the chart, and is it falsifiable?), plus logging so that runs can be reproduced.
  • Across the dataset: splitting the train and test sets by patient rather than by question, so prompt tuning never sees a held-out patient’s record.

“Three ways a synthetic perturbation can fail to be a useful test”

Once they observed these, what did they do to avoid them? Prompt adjustments?

From the post: Each failure mode got a different fix, applied at a different stage:

  • Accidentally correct (for example, they swapped in azathioprine, but the patient really did try it): “Detect & exclude from metrics,” which is a check after generation.
  • Unfalsifiable (for example, “all alternative therapies were tried”): “Require specific, falsifiable claims,” which is a constraint on generation.
  • Administrative noise (attestations): “Filter out before perturbing,” which limits which questions are eligible.

Inference: The post doesn’t say how these were implemented. The likely mix is prompt instructions for falsifiability, a question-type filter for attestations, and some kind of validator for accidentally correct cases. That validator has a trap. If it relies on the judge’s own “Supported” verdict to decide that a perturbation was accidentally correct, it will also throw out real misses and inflate the detection rate. The check has to be independent of the judge, for example a lookup in the structured medication history, or human review of every excluded case.

“In production, real errors often cascade through related questions”

What does this mean?

Inference: A PA form asks many related questions about one patient, and the generator answers all of them from the same reading of the chart. If that reading is wrong at one point, the error spreads to every answer that depends on it. For example, if the generator takes the diagnosis to be ulcerative colitis, it may also list therapies relevant to that disease, report a different severity score, and argue medical necessity for the wrong condition. So real errors are correlated across questions, while their perturbations change one answer at a time.

This matters for the consistency signal. In a full cascade, all the answers could be wrong in a way that is consistent with each other. The cross-answer check then wouldn’t fire, and only checking against the chart would catch the error. Single-answer perturbations may therefore make consistency-based detection look more useful than it really is.

“the judge flags a different, unperturbed answer because it conflicts with the perturbed one”

Is their judge actually looking for inconsistency across questions or is this a hypothetical?

From the post: It’s described as observed behavior, not a hypothetical: “The judge flags two things,” and the second is “a behavior we hadn’t explicitly designed the judge to produce.” They also say they have to separate direct detection from consistency-based detection in their evaluation, which would only be necessary if it actually happens in their runs. The sentence you quoted is how they define that second category. They don’t report how often it happens, though, and the interactive demo is scripted.

Inference: Since they didn’t design it, the judge isn’t explicitly checking for consistency across questions; the behavior comes from somewhere else. A plausible mechanism is decomposition. For Q2 (“What clinical features support the diagnosis?”), decomposition that draws on question context would produce an implied claim like “these features support a diagnosis of ulcerative colitis,” taking the diagnosis from Q1’s answer. The chart contradicts that claim, so Q2 gets flagged. If that’s right, consistency detection only works when one question explicitly refers to another answer.

“One is mining production data for confirmed errors”

How is an error confirmed?

Inference: Possible sources of confirmation, from most to least direct:

  • Reviewer outcome: a human checks a flagged answer against the chart and corrects it.
  • Provider correction: the prescribing office amends an answer after seeing the form.
  • Payer response: a denial or a request for more information that points at a specific answer. Denial reasons are often vague, so linking them to an answer takes manual work.
  • Audit sampling: running the gold-standard annotation process on random production cases, including ones the judge passed.

The first three mostly surface errors someone already noticed, and the reviewer route only confirms errors that the judge itself flagged. Errors the judge misses won’t show up unless something downstream catches them. So mining confirmed errors under-represents exactly the misses the evaluation most needs, and random audits are the way to correct for that.

“Another is building an error taxonomy from real cases so that perturbations cover the actual distribution of mistakes rather than the ones a model finds easy to invent”

How is this different from mining production data in the previous sentence?

Inference: They work at different levels:

  • Mining gives you individual real errors. These can seed the perturbation model directly, for example as few-shot examples, which makes each perturbation look more like a real mistake.
  • A taxonomy classifies those errors and records how often each type occurs. Example types include a wrong date range, an outdated lab value, past history stated as current, a wrong drug in the prior-therapy list, or a flipped negation. The taxonomy tells the generator how many of each type to produce, so the test set matches the real mix. Otherwise it would over-represent errors that are easy for a model to invent, such as swapping a diagnosis.

Mining makes each example more realistic, and the taxonomy controls which kinds of errors appear and in what proportions. A taxonomy also lets you report detection rate per error type, which shows where the judge is weak. You build the taxonomy from mined errors, so mining comes first.

“where the perturbation model is trained to produce errors the judge misses”

How would you actually do this? Prompt iteration, fine-tuning or something else?

Inference: Three options, from cheapest to most involved:

  • Search without training: generate many candidates, keep the ones the judge misses, and feed them back as few-shot examples in the next round. No model weights change.
  • Preference fine-tuning (for example DPO): train on pairs in which the missed perturbation is preferred over the caught one.
  • Reinforcement learning: reward the generator whenever the judge misses its error.

Every option needs a validity check that doesn’t depend on the judge. Otherwise the generator learns the easiest way to beat the judge, which is to produce accidentally correct or unfalsifiable perturbations, the same failure modes the post documented. A perturbation should count as a success only if it meets three conditions. It must be verifiably wrong against the chart (checked by structured lookup, a separate checker that sees the gold answer, or human spot checks). It must be clinically plausible. And the judge must miss it. Restricting generation to the error types in the taxonomy (Q26) also keeps the errors realistic.

What could “production monitoring” actually look like for their use-case?

Inference: Grouped by what each signal tells you:

  • Over-flagging: the share of flags that reviewers overturn, broken down by drug, payer and question type.
  • Missed errors: payer denials or information requests on answers the judge passed, plus random audits of unflagged answers.
  • Drift: changes over time in the flag rate and in the mix of verdicts. A rising share of Not Addressed claims, for example, can mean that retrieval broke or that a new EHR integration is sending incomplete records.
  • Retrieval health: the share of claims for which nothing was retrieved.
  • Patient impact: time from prescription to submission for flagged versus unflagged cases, which shows what over-flagging actually costs.

A useful addition is canary testing: regularly running known perturbed forms through the live pipeline to confirm that detection hasn’t dropped after a model or prompt change.

How would you improve this system if you were building it today?

Grouped by component:

  • The judge: make the cross-question consistency check a deliberate step instead of a side effect, so it can be measured and can catch full cascades (Q23). Route claims that payer criteria depend on separately, so a Not Addressed verdict on one of them goes to a documentation queue (Q17).
  • The flagging threshold: have the judge output a confidence for each claim, and set the threshold from the relative cost of a missed error versus a false flag, instead of using a fixed worst-verdict rule.
  • Perturbations: generate them from an error taxonomy (Q26), include errors that cascade across several answers, and check for accidentally correct cases independently of the judge (Q22).
  • Evaluation reporting: report detection rate by error type with confidence intervals. Also keep a small held-out set of real confirmed production errors as the final test, since synthetic detection rates probably overstate performance on real errors.
  • Monitoring: add random audits of unflagged answers, which are the only unbiased way to estimate how many errors the judge misses in production (Q25).
2025 Blogpost original

Building and Deploying Production‑Grade AI Agents: Cresta’s End‑to‑End Approach

My notes

Some key insights/new things I learned

  • Two useful components for the researching phase prior to agent development
    • Insights (Topic Discovery): provides a bird’s-eye view of conversation topics and outcomes
    • Automation Discovery: reconstructs representative dialogue flows from the transcripts
  • Simulated Visitors
    • Generated from de-identified real conversations
    • Deliberately enable imperfect communication and understanding to resemble real-life

What I didn’t understand or am unsure about

Question

Answer

“In the case of AI agents, Cresta suggests how to modify the agent’s prompt or other parts of its configuration and then runs tests to validate that the agent’s performance improved as a result.“

Does “Cresta” here refer to a Cresta employee or a Cresta system? In other words, is this done manually or automatically?

[Source] The post doesn't say. The sentence is built as a parallel to the preceding one about human agents, where Cresta the platform delivers hints and coaching, so the grammar points toward the system.

[Inference] The two verbs are doing different amounts of work. "Runs tests" is plainly automated — that's the evaluator toolkit and simulation suite described later. "Suggests how to modify" is the softer claim, and most likely means the platform surfaces a diagnosis (this turn failed this behavior, here's why) while a person decides the actual edit. The post's own Lessons Learned section supports that reading, since its entire argument is that humans stay in the loop on targeted decisions. So the honest summary is automated detection and automated validation, with a human-authored change in between. "Suggests" is the kind of verb that lets a vendor describe either arrangement.

Are Cresta’s agents primarily text-based or voice?

[Source] This post is modality-agnostic. It refers to "chat and call transcripts" and mentions both inbound and outbound conversations.

[Search] Voice is clearly the emphasis across the rest of their writing. Their Voice Virtual Agent announcement says voice is "the top used channel at over 70% of interactions" and criticizes competitors for "focusing exclusively on chat," and their latency engineering post is entirely about the voice pipeline — speech recognition, voice activity detection, speech synthesis.

[Inference] Both channels are supported, but the engineering problems they choose to write about are voice problems. The "interruptions" metric listed in this post only makes sense for voice at all. I'd read the post as voice-first with chat along for the ride.

“Cresta fixes this cold start problem by grounding the design in historical data”

Doesn’t cold start imply no historical data?

[Inference] You're right that the term is being stretched. Cold start conventionally means having no data about a new user or item. Here the cold start belongs to the agent, not the enterprise: the new AI agent has no interaction history of its own, while the enterprise has millions of transcripts. So what they're describing is a transfer-learning move — bootstrap the new agent from the behavior of the incumbent human agents. That's a reasonable idea, but calling it "fixing cold start" is marketing language; a stricter description would be a warm start from a related distribution. And the distributions genuinely differ: customers ask different things of a bot than of a person, and they behave differently once they know which one they're talking to.

What would you do if you actually didn’t have any historical data and wanted to launch something like this?

[Inference] Roughly in order of cost, I'd do five things. First, mine the written artifacts that exist even without transcripts — standard operating procedures, help-center articles, canned agent responses, past ticket subject lines, and abandoned queries typed into a chat widget. Those give you intents without needing dialogue. Second, have subject-matter experts record twenty to fifty think-aloud role-plays, which becomes the seed corpus your personas are drawn from. Third, generate synthetic conversations from the procedure documents, but validate them against the role-plays so you aren't only testing the agent against your own assumptions. Fourth, launch deliberately narrow, or run the agent in shadow mode alongside human agents so it produces responses that get scored but never sent. Fifth, instrument heavily, because the first few thousand real conversations are the actual asset and you want the relabeling loop to be cheap from day one.

The thing to watch is that the first three all encode the same prior. Only real traffic breaks it.

“finds the major clusters of customer issues”

How?

[Search] Their Insights announcement is more specific: Topic Discovery uses "unsupervised LLM clustering on conversation reasons," and generates "an exhaustive taxonomy and up to three levels of subcategories" without teams having to predefine categories.

[Inference] In practice that reads as a two-stage pipeline. An LLM first extracts a short conversation reason from each transcript, compressing a long call into a normalized phrase. Clustering then runs over those compressed reasons rather than over raw transcripts — most likely embeddings plus a clustering algorithm, with an LLM naming clusters and merging them hierarchically into the taxonomy. Clustering on extracted reasons instead of full text is the important design decision, because small talk, greetings, and hold messages would otherwise dominate the similarity signal and you'd end up clustering conversations by style rather than by problem.

Why do so many of their cluster examples sound very similar to each other?

[Inference] Two explanations, and they have different implications. If the screenshot is showing sibling leaves beneath one parent node, then similarity is expected and fine — "billing, unexpected charge" and "billing, duplicate charge" are supposed to differ by one discriminating detail, and the three-level taxonomy makes that likely. The less comfortable explanation is under-merging, which is a real failure mode of LLM-driven taxonomy induction. The model finds distinctions that are linguistically genuine but operationally meaningless, because nothing in the clustering objective penalizes redundancy. The test I'd apply: if two clusters would route to the same conversation flow and trigger the same tool calls, the taxonomy is over-split for the purpose it's being used for. Your suspicion is worth holding onto.

“how human agents resolved them”

Do we know for sure that the human agents' approaches were correct? How?

[Inference] "Proven" is doing unearned work there. Transcripts record what humans did, not what was correct. There are three ways to get closer to correctness, none of them free. You can filter by outcome rather than by behavior, keeping only conversations with a good resolution signal such as no callback within some window, positive satisfaction score, or no subsequent refund reversal — which turns plain imitation into outcome-weighted imitation. You can stratify by agent and learn only from the top quartile on the relevant metric, which is the standard contact-center move and something Cresta's quality management product already scores. Or you can treat the mined flow as a hypothesis rather than ground truth and put it in front of subject-matter experts before it becomes the spec. There's also a survivorship risk worth naming: shortcuts agents use to hit handle-time targets look indistinguishable from best practice in a transcript.

“Insights surfaces real customer phrasings and look-alike intents”

What does this mean?

[Inference] They're two separate things. Real customer phrasings means the utterances people actually produce, as opposed to how the business names the intent internally — customers say "I got charged twice," while the internal taxonomy says "duplicate transaction dispute," and it's the former you need in the router's examples. Look-alike intents means pairs that sit close together in surface form but need completely different handling. "I want to cancel my order" and "I want to cancel my subscription" share most of their words and then diverge into different flows touching different systems. Surfacing those pairs during scoping tells you in advance where your router will confuse itself, so you can add a disambiguation turn or supply discriminating examples before the confusion shows up in production traffic.

“we can explicitly define what the AI agent will handle and what falls out of scope”

What are realistic examples of things that would fall out of scope?

[Inference] Exclusions tend to fall into four groups. There's regulated or liability-bearing work — debt collection promises, insurance coverage determinations, medical or legal advice, anything that creates a binding commitment on the company's behalf. There's high blast radius — refunds above a threshold, account closures, changing a payment method, anything hard to reverse. There's emotionally loaded contact — bereavement, fraud victims, threats of self-harm, and retention saves where human judgment carries the conversation. And there's the long tail, meaning intents that appear a handful of times a month where the cost of building and maintaining a flow exceeds the containment you'd gain. By raw count the long tail is the largest category. The first two are the ones that get written into the contract.

How is “automation readiness” determined?

[Source] This post lists it as something Insights quantifies, with no method given.

[Search] Their dedicated Automation Discovery post is more forthcoming: "The score is driven by structural and operational signals observed in the data, including the frequency and nature of deviations, as well as the level of inferred integration complexity."

[Inference] Unpacking those two signals, deviation frequency is a proxy for how well-structured an intent is — if most conversations follow one dominant path, a state machine can capture it, and if they scatter, the intent needs judgment. Integration complexity is a proxy for build cost, meaning how many backend systems the flow has to touch. So readiness is roughly path regularity divided by integration cost, weighted by volume. Worth noticing that this is a build-feasibility score, not a risk score. Nothing in it speaks to whether automating a given intent is safe, and the post's framing quietly blends the two axes.

“Automation Discovery reconstructs representative dialogue flows from the transcripts”

If this tool is looking at a very large number of transcripts (presumably not fitting in a single context window), how is it generating the desired information correctly?

[Source] Not addressed here. Their Automation Discovery post says only that they use "LLMs together with deterministic analysis."

[Inference] Nobody puts a million transcripts into one context, so the shape is almost certainly a four-stage pipeline. First, per-transcript extraction, running one model call per conversation in parallel, compressing each into a structured trace: the reason, an ordered list of steps, tool-like events, and the outcome. This is the only stage that sees raw text and it's parallel. Second, cluster the traces rather than the text. Third, align the step sequences within each cluster, which is closer to process mining than to language modeling and can be done deterministically with frequent-sequence mining or transition counting. Fourth, bring an LLM back in only at the end, to name states and write descriptions over an aggregate that now fits comfortably in context. Their phrase "deterministic analysis" most plausibly points at the third stage. The quality risk sits in the first: whatever schema the extractor uses silently bounds what the final state machine is capable of expressing.

“State machine skeleton: It identifies the dominant paths customers take to reach resolution”

Per intent, or more general?

[Inference] Per intent, almost certainly. A single global state machine spanning all customer issues would be close to useless, because the states for rescheduling an appointment and disputing a charge share nothing except authentication. The pipeline ordering in the post supports this: Insights clusters issues first, then Automation Discovery reconstructs flows, so the natural unit is one flow per cluster, with shared preamble states like greeting and identity verification and shared exit states like escalation factored out as common subgraphs. Their decentralized-architecture post is consistent with this, since it names sub-agents for "authentication, policy verification, troubleshooting workflows, payment handling" — which is exactly a shared-preamble-plus-per-intent-body decomposition.

Is a human the primary consumer of the output of the Automation Discovery tool to guide the development of the agent? Or is it being directly fed into a system somewhere?

[Search] Both, according to their Automation Discovery post. It supports human inspection, and it "can export a structured draft prompt derived from the workflow, serving as a scaffold for AI agents."

[Inference] "Draft" and "scaffold" are the operative words — this is a human-reviewed handoff, not a closed loop. That's the right call given what the artifact is. The mined flow encodes whatever the human agents happened to do, so pushing it straight into production would automate their mistakes at scale. The post you're reading agrees implicitly by calling the scoping output a "v1 spec," which presupposes somebody reviews it.

“Tool catalog: identifies critical data dependencies and system touchpoints”

How can you get this information just from transcripts?

[Search] Their Automation Discovery post answers this directly: "When agents say things like 'let me check your order,' 'I'll submit a claim,' or 'I need to verify eligibility,' those patterns often indicate interaction with backend systems."

[Inference] So it's linguistic evidence that a system call happened, not observation of the call itself. Two further signals are usually available and they don't mention either: long silences or explicit hold events in a call mark a lookup, and the shape of the data the agent reads back — an order number, a balance, a ship date — reveals both which system was queried and roughly what it returned. The limitation is real, though. This gives you a hypothesis list of integrations with no guarantee of completeness, no API contract, and no visibility into any system the agent used without narrating it. Screen recordings or CRM audit logs would be the ground truth; transcripts are the cheap proxy for it.

How is “Tool use and data needs” different from “Tool catalog”?

[Source] The post lists them as separate bullets and the distinction it draws is thin. Tool catalog identifies "critical data dependencies and system touchpoints" and enables function schema specification; tool use and data needs catalogues required internal systems and data such as CRM lookups and order adjustments, and informs function call specifications.

[Inference] As written these overlap heavily, which by a MECE standard is a defect in the post rather than a distinction you're missing. The most charitable reading is that the catalog is the inventory — which systems exist and get touched — while tool use and data needs is the contract, meaning what each call requires as arguments and what it returns. Inventory versus signature.

“that might merit separate specialized agents”

Is the appeal of separate agents a clean context window or different capabilities (whether through a different model, harness or skills)?

[Source] This post doesn't say.

[Search] Their decentralized-architecture post makes the case, and the reasons it gives are mostly neither of your two options. It cites avoiding a single point of failure, since if a central orchestrator "falters due to an unexpected scenario, a model update, or a spike in call volume, the entire experience may degrade"; avoiding errors that "accumulate across long conversations"; and latency, because "every decision and turn reroutes through the orchestrator." It does also describe keeping "the AI focused on the immediate step," which is your clean-context argument.

[Inference] In practice I think prompt focus is the dominant real benefit, so your first option. A sub-agent's prompt contains only the rules for its own task, which raises instruction-following accuracy and cuts token cost. Independent testability is second, and this post names it outright — "each skill can be built and tested independently." Running different models per sub-agent is easy to do and occasionally worth it, a small fast model for authentication and a stronger one for troubleshooting, but it isn't the headline reason.

“and a world state the agent must navigate: authentication, entitlements, status, partial data, the messy stuff”

Why is this in the simulated visitor spec as opposed to the test spec? Does their version of visitor both the simulated user and the task?

[Source] Yes, it bundles them. The simulated visitor spec as described carries four things: a persona with tone and vocabulary, a goal defining "what 'done' means," the world state, and behavior policies covering how the person asks, clarifies, and pushes back. Separately the post says "we know 'what should happen' in each scenario," so the expected outcome travels with the scenario too.

[Inference] Your instinct that these are separable is right, and separating them would be the better design. Persona, world state, and goal form a cross-product you'd want to sample along independently, so you could hold a persona fixed while varying account state, or the reverse. Bundling them into a single artifact means every combination is a hand-built object and your coverage is limited to whatever someone thought to write down. The likely reason it's bundled is how they're generated: if each visitor is derived from one real conversation, the persona and the world state come from the same source record and arrive already entangled.

“Because these simulated visitors are generated from de-identified real conversations”

How generated?

[Source] No method given beyond the de-identification.

[Inference] The plausible pipeline is to take a real transcript, strip personally identifying information using named-entity redaction plus format-based scrubbing for things like card and account numbers, then have an LLM read the cleaned transcript and emit the structured spec. It would infer the persona from how the customer wrote or spoke, the goal from what they were trying to achieve, the world state from facts established during the conversation — were they authenticated, did they have an active policy, was the order partially shipped — and the behavior policies from observed conduct such as pushing back, wandering off-topic, or misunderstanding a rule. Then you discard the original transcript and keep only the spec, which is what makes the artifact privacy-safe: it's a description of a customer rather than a copy of one. Whether they check the generated spec back against the original for fidelity is unstated, and that's the thing I'd most want to know.

“simulate full multi-turn conversations between each virtual customer and the AI agent”

  1. How to sandbox to ensure no cross-context contamination?
  2. Different models or same? Does it matter?

[Inference] (a) The contamination risks are distinct and each needs its own control. Within a single run, the visitor model must not see the agent's system prompt, tools, or internal state — only its own spec plus the dialogue so far — which means two separate processes with separate contexts, not two personas sharing one context. Across runs, you need no shared memory, fresh conversation identifiers, and a reset of the mock backend between runs, because otherwise run N's writes, a payment recorded or a ticket opened, silently change run N+1's world state. The agent's own tools have to hit mocks or a seeded test database rather than production, and this is the one that catches people out, since a tool that looks read-only often still writes an audit record. Finally the grader shouldn't know which configuration produced a transcript, or you introduce scoring bias precisely when you're diffing prompts against each other.

(b) Different models, and yes it matters. Using the same model on both sides is the classic evaluation trap: the visitor and the agent share blind spots and failure modes, so the agent scores better than it deserves and the visitor never generates the confusions a different model — or a real person — would produce. Running a different family for the visitor, and ideally a third for the judge, buys you independence. The cost is that your simulated population now inherits whatever quirks the visitor model has, which is exactly why the pool still needs to be grounded in real transcripts.

“The result of each simulation run is a scored transcript with precise reasons for any failures and artifacts users can diff across prompts, tools, or models.”

Same scoring methodology as online and human agents?

[Source] Partly, and the post is explicit about the shared part. It says "a lot of AI agent requirements map exactly to behaviors and are the same for human and AI agents," and that customers running both "can often share the same behavior definitions between the two." That's about Opera behavior definitions applied in production.

[Inference] But simulation scoring has something production scoring structurally cannot have, which is a known expected outcome — the post's "we know 'what should happen' in each scenario." So simulation can score outcome correctness by direct comparison, while production can only score behavior adherence by judgment, because nobody knows the counterfactual for a live call. My read is a two-layer scheme: shared behavior evaluators, the Opera definitions, applied everywhere, plus an outcome-comparison layer that exists only offline. That asymmetry is worth holding onto, because a high simulation pass rate and good production behavior scores are not measuring the same thing and can move independently.

“Because the simulator can spin up hundreds of variations of a scenario, it’s great for catching edge cases that never surfaced in a low-volume manual test.”

How to encourage diversity and exploration to uncover more edge cases?

[Inference] Naive sampling from one model collapses toward the mode, and you end up with hundreds of near-identical polite customers. The things that actually work: factorize and grid rather than sample, meaning enumerate the axes explicitly — authentication state, data completeness, emotional register, language proficiency, interruption behavior — and cover the cross-product instead of asking a model for a hundred varied customers. Seed from the tail of real data by sampling source transcripts inversely to cluster frequency, so rare intents are over-represented in the visitor pool. Generate adversarially by conditioning the visitor on a known failure mode, such as insisting on a policy exception, giving the account number with one digit wrong, or changing goal mid-conversation. Apply novelty pressure by embedding generated scenarios and rejecting any that sit too close to an existing pool member, which is cheap and surprisingly effective. And mine production for what you missed, which is what they describe doing post-launch and is the only genuine source of unknown unknowns.

The measurable version of all this is coverage over the state machine's transitions: which edges has the pool never exercised?

“after go-live, we also keep expanding the pool with new conversation patterns as real users interact with the agent”

Automatically or manually? How?

[Source] Unstated. The full sentence continues "so we can continually run regression tests."

[Inference] The mining step is almost certainly automatic, since it's the same visitor-generation pipeline from question 18 pointed at production transcripts instead of historical ones. The interesting decision is candidate selection, and the natural signals are conversations where the agent escalated, where the customer repeated themselves or expressed frustration, where an evaluator flagged a behavior violation, and where the conversation reason falls outside the existing taxonomy.

[Search] That last one is exactly what their Trends feature detects — their Insights post describes Real-Time Trends surfacing "emerging phrases, unusual spikes, and cross-cutting patterns," which is a natural feed into this.

[Inference] Promotion into the regression suite should be manual, or at minimum reviewed. An automatically added test encodes whatever the agent did as the expected behavior unless a human supplies the expected outcome, and the post's own emphasis on knowing "what should happen" implies someone has to.

“encapsulates everything the agent needs: its prompts, decision logic”

What is “decision logic”?

[Source] The post doesn't define it. It lists the term alongside prompts, tool integrations, and guardrails, and elsewhere mentions adjusting a "workflow stage" as a thing you can do to a config.

[Search] Their decentralized-architecture post is more concrete, describing "deterministic state management" that keeps "track of exactly where a customer is in a process, triggering the right actions at the right moment," combined with dynamic prompting that keeps "the AI focused on the immediate step."

[Inference] So decision logic is the non-LLM part of control flow — the state machine from the scoping phase, rendered as configuration. Which stage the conversation is in, which transitions are legal from here, which tool fires on entering a state, when control hands off to another sub-agent, and when the conversation escalates. The language model handles what gets said within a state; the decision logic constrains which states are reachable at all. That's the hybrid architecture they advertise, and it's also why a config bundle is diffable in a way a single prompt isn't: state transitions are structured data, prompt prose is not.

“Under the hood, the agent’s reasoning and capabilities are extended by Cresta’s AI Agent Framework which is a declarative framework that lets developers register backend functions”

What is “a declarative framework”?

[Inference] Declarative means you describe what something is and the framework works out how to execute it, as opposed to imperative, where you write the control flow yourself. Concretely here, a developer writes an ordinary function and annotates it with a schema — a name, a description, typed parameters, a return type — and the framework handles everything downstream: exposing it to the model as a callable tool, validating the model's arguments against the schema, invoking the function, and feeding the result back into the conversation. The developer never writes the "if the model requested this tool, parse its arguments and dispatch" plumbing. Python decorators are the usual surface syntax for this. The real payoff is that the tool's declaration and its implementation are the same artifact and therefore can't drift apart — which matters, because "schema mismatches" is one of the failure types the post names in its optimization loop.

“The framework is optimized for ultra-low latency”

How?

[Source] The post asserts it without method, tying it only to "preserving a smooth real-time conversation."

[Search] Cresta has a whole post on this. Their stated budget is that pauses beyond about 300 milliseconds "can feel unnatural" and anything past roughly 1.5 seconds "can rapidly degrade the experience." Component budgets they cite: audio preprocessing 25 to 50 ms, speech recognition 200 to 300 ms, turn detection at least 600 ms, language model time-to-first-token anywhere from 250 ms to over a second, and speech synthesis first byte 100 to 500 ms. The named techniques are streaming APIs with no DNS lookup in the critical path, reused connections to the model, WebRTC transport which they say "can reduce latency by up to 300ms," guardrail model calls issued concurrently with the main call rather than after it, speculative triggering meaning "starting the LLM call before the user fully stops speaking," hedging meaning "launching multiple LLM calls in parallel and using whichever returns first," and smaller models in the live loop since "reasoning models generally can't be used within the live response loop."

[Search] On tool calls specifically, which is what the framework sentence is really about, their sharpest observation is that if you look something up before the model can say anything at all, "the 1 LLM first-token latency becomes 1 LLM latency + 1 LLM first-token latency." Their mitigations are to fire predictable lookups concurrently at the start of the conversation, emit filler or wait messages for calls under about ten seconds, and run anything longer asynchronously.

[Inference] That last cluster is the actual content behind "optimized for ultra-low latency." The framework's job is to make a registered function's round trip overlap with generation rather than block it.

“Cresta’s framework supports this composition, so each skill can be built and tested independently”

What does skill mean here?

[Source] The post uses skill more or less interchangeably with specialized sub-agent — the preceding sentence is about "specialized sub-agents for distinct tasks." It also uses "micro-skills → agents" back in the scoping section, describing task decomposition that informs multi-agent architecture.

[Search] Their decentralized post names the actual units: "authentication, policy verification, troubleshooting workflows, payment handling, and more."

[Inference] So a skill here is a bounded task with its own prompt, its own tools, and its own entry and exit conditions — one coherent chunk of the state machine, packaged so it can be exercised in isolation. Worth noting this is not the Anthropic sense of "skill" as a loadable instruction bundle; it's much closer to a microservice with a prompt attached. A good test for whether something deserves to be its own skill is whether you can write a pass/fail condition for it without referring to the rest of the conversation.

“then unified in an orchestrated flow”

By an orchestrator agent?

[Source] The post says "orchestrated flow" and leaves the mechanism unspecified.

[Search] Their architecture post argues explicitly against a central orchestrator agent. It says a central orchestrator "can also create a single point of failure," that errors "accumulate across long conversations," and that centralized designs "introduce extra steps, since every decision and turn reroutes through the orchestrator." Their alternative is "a network of specialized subagents" that "handle their areas of expertise directly and in parallel," coordinated by deterministic state management.

[Inference] So "orchestrated" here means orchestrated by configuration, not by a supervising model. The state machine decides which skill is active and when control transfers, and there's no LLM sitting above the sub-agents making a routing decision every turn. That's a substantive architectural claim and a defensible one for voice, since it removes a model call from the critical path and makes handoffs deterministic and independently testable. The tradeoff is rigidity: anything the state machine didn't anticipate has no path through the graph and falls through to the escalation edge.

“The optimization loop is simple: Tweak the agent config or remote functions, run the test set, and fix whatever shows up in red (bad routes, schema mismatches, guardrail violation etc)”

Overfitting concerns?

[Source] Not addressed. The post presents the loop approvingly and notes that every run is archived for comparison against the last green one.

[Inference] Yes, and I'd call this the sharpest methodological gap in the post. Fixing whatever shows up red and then rerunning the same test set is gradient descent on the test set with a human as the optimizer. Four specific risks follow. You get test-set memorization, where prompt edits patch individual failing scenarios rather than the underlying behavior, which shows up as a pass rate that climbs while production quality doesn't move. Nothing in the description mentions a held-out split the config is never tuned against, which is the minimum defense. With hundreds of stochastic scenarios and repeated runs you get multiple-comparison drift, where some red-to-green flips are noise and chasing them adds prompt text that constrains nothing real. And prompt bloat is the visible symptom of all of it: each patch appends a rule, the prompt grows, instruction-following degrades, earlier rules get quietly ignored, and the fix for failure N breaks passing case M.

What I'd expect in a mature version: a frozen holdout set, freshly generated scenarios each cycle, several seeds per scenario with a significance test on the difference rather than eyeballing red versus green, and a regression gate that reruns previously-passing cases. Their post-launch mining partly addresses this by continuously refreshing the pool, which is the right instinct — but it only works if new scenarios are being added faster than they're being tuned against.

“If it clears the requirements (pass rate, no criticals, tool-accuracy targets)”

  1. What is pass rate?
  2. What are tool-accuracy targets and why do we care about them?

[Inference] (a) Pass rate is the fraction of simulated scenarios where the agent reached the expected outcome, with the post's "we know 'what should happen'" providing the comparison basis. It's a conversation-level metric, which their companion post distinguishes from turn-level checks. Two things need pinning down before a pass rate means anything across versions: what counts as a pass, meaning exact outcome match versus evaluator-judged acceptability, and how the scenario pool is weighted, since a pass rate computed over a pool that over-samples easy intents isn't comparable to one computed over a harder pool.

(b) Tool accuracy is whether the agent called the right backend function, at the right moment, with correctly-formed arguments. It decomposes into at least four failure modes worth tracking separately: the wrong tool was selected, the right tool was called at the wrong point in the flow, the arguments were malformed or hallucinated — the "schema mismatches" the post names — or no tool was called when one was required.

It earns its own gate for three reasons. It's where the irreversible damage lives, since a clumsy sentence is recoverable and a refund issued to the wrong account is not. It's the one dimension where the agent touches systems of record, so its errors escape the conversation and persist. And unlike conversational quality it's cheaply and objectively checkable, because you know what the correct call was, which makes it a good hard gate rather than a soft score.

“Once an AI agent is live and serving production traffic, we monitor its performance”

Who is “we” here? Cresta or the customer?

[Source] Ambiguous, and the referent shifts within the section. The sentence continues "much like we monitor human agent conversations," and human-agent monitoring is done by the customer's quality team using Cresta's tools. But the following sentences say "we monitor all the metrics we care about tracking," and then describe "making it easy for our customers to align the system," which puts customers in the third person.

[Inference] Most likely both, with the weight on the customer. Everything described — Dashboard Builder, Opera, AI Analyst — is a self-serve product the customer operates, and the entire framing of Opera is that customers define their own behaviors. Cresta presumably also watches closely during deployment and ramp. The slippage in "we" is worth noticing because it conceals an operating-model question: is post-launch tuning a service Cresta performs, or a product the customer runs? The post reads as the latter, sold with the former during onboarding.

“This includes metrics common to AI agents, or specific to certain types of AI agents. For example: Interruptions”

What are interruptions?

[Inference] An interruption is the customer speaking while the agent is still speaking — barge-in. It gets tracked because it's a cheap, automatically detectable proxy for several distinct problems at once: the agent is talking too long, the turn-detection threshold is mis-tuned so the agent starts before the customer has finished, the response was wrong and the customer is cutting in to correct it, or the customer is simply frustrated. A rising interruption rate is an early warning that requires nobody to read a transcript. Its weakness is exactly that ambiguity — it won't tell you which of those four is happening, so it's a trigger for investigation rather than a diagnosis.

“Trends and Anomalies page automatically detects shifts in the distributions of conversation reasons”

Is this mechanistically different from Insights?

[Search] Their Insights post draws the distinction directly. Topic Discovery provides "durable structure: a trusted view of what customer conversations are about," while Real-Time Trends provides "the real-time signal: visibility into what is changing right now, including events no taxonomy was built to catch."

[Inference] So they're layered rather than alternative, and the mechanical difference is the unit of work. Topic Discovery is a batch clustering job that induces a taxonomy: expensive, run periodically, output is a stable label set. Trends and Anomalies is a streaming statistical job over labels that have already been assigned: cheap, run continuously, output is a change signal on distributions you already defined. The phrase about "events no taxonomy was built to catch" implies a second component that isn't just distribution monitoring — emerging-phrase detection over unlabeled text, which catches the thing that doesn't have a bucket yet. That's the piece that genuinely overlaps with clustering, and is presumably a lighter-weight version of it run over a recent window.

“Show them a specific conversation turn where the agent did something questionable”

How do you know in advance if an agent did something questionable without showing it to a human?

[Source] The post doesn't address the selection problem, which is the load-bearing gap in its Lessons Learned section. Their companion evaluator post doesn't either — it covers how to calibrate evaluators, including a preference for binary over numeric scales because "numerical scales are inherently more subjective than binary classifiers," but not how candidates get chosen for review in the first place.

[Inference] This is the classic active-learning problem, and the key realization is that you don't need to know the right answer to select well. You need a cheap signal that correlates with being wrong. The usable ones, roughly in order of yield: evaluator disagreement, where you run two or more judges with different prompts or models and route every disagreement to a human; evaluator uncertainty, where the judge's score sits near its decision boundary or its self-consistency across sampled runs is low; hard rule violations, which are detectable deterministically and always worth surfacing — a promise made, a verification step skipped, a tool called before authentication; outcome signals such as escalations, customer repetition, negative sentiment, or a callback within 24 hours, which need no judgment at all because the conversation tells you it went badly; and novelty, meaning turns far from anything in the labeled set, where the evaluator is least trustworthy.

The honest framing is that you can't know a turn is questionable, only that it's uncertain or anomalous. What that buys you is throughput: it converts "read a thousand transcripts" into "answer fifty yes-or-no questions," which is precisely the fatigue problem the post's lesson is about. The failure mode to watch is that uncertainty sampling systematically misses confident-but-wrong turns, so a small random audit sample has to run alongside it.

“Every clarification from a human should either adjust the agent or its evaluation guidelines going forward”

Adjust in what way? Prompt?

[Source] Verbatim, and the sentence continues "so we don't ask the same question twice." The mechanism is not specified.

[Inference] The sentence names two destinations, and they take different kinds of edit. Adjusting the agent means changing the config bundle, and the post's own framing says a config is much more than a prompt — so the right edit depends on the failure. A wrong tool call is a schema or tool-description fix, not a prompt fix. A missing empathy line is a prompt fix. An illegal path through the conversation is a state-machine fix. A hard compliance rule shouldn't go in the prompt at all; it belongs in deterministic control flow. Adjusting the evaluation guidelines is the other half and it changes the judge rather than the agent: when a human says "no, that response was actually fine," the agent was right and the evaluator was wrong, so the rubric gets amended.

[Search] That second half is what their evaluator post is about — aligning stakeholders on what a behavior like "Agent Assumes Payment" means in a given context, where "each misalignment creates an opportunity to refine the guideline."

[Inference] The requirement hiding inside "so we don't ask the same question twice" is that each clarification has to be stored as a durable artifact with provenance — a rubric clause or a config line — rather than applied once and forgotten. Otherwise the same ambiguity resurfaces with the next reviewer.

“The key is to have a process that rapidly discovers and fixes failures when they occur.”

What part of their pipeline actually provides rapid feedback?

[Source] The post answers this in the very next sentence: "That's why the simulated testing and the post-launch analytics are so vital." No turnaround times appear anywhere in the post.

[Inference] Discovery and fixing have different bottlenecks and it's worth separating them. Discovery is the genuinely fast part, and it's the post-launch layer that delivers it: Trends and Anomalies flags distribution shifts without anyone asking, the Opera behavior evaluators score every conversation rather than a sample, and AI Analyst answers an ad-hoc question in minutes instead of requiring an analyst to pull and read transcripts.

Fixing is where the clock slows. The loop is edit config, run the simulation suite, compare against the last green run, lock the version, push. Simulation is the part that parallelizes, but the post never says how long a full suite takes, and that number is what actually sets cycle time.

[Inference] What's conspicuously absent is anything between "detected in production" and "fixed in the next version" — no hot-fix path, no feature flag, no way to disable one misbehaving skill while the real fix is built. Config versioning and rollback are the closest thing mentioned, but rollback reverts everything rather than the one broken flow. For a voice agent carrying live traffic, that gap matters more than the length of a test run.

“We learned to prioritize requirements by criticality.”

How does this prioritization actually manifest?

[Inference] It should show up in four concrete places. In the release gate, where the criteria quoted in Q29 are "pass rate, no criticals, tool-accuracy targets" — "no criticals" is a hard blocker while pass rate is a tunable threshold, so criticality manifests as a different kind of gate, not a heavier weight. In where the behavior is enforced, since critical requirements shouldn't live in prompt text at all: you cannot reach the account-change state without passing through verification, and that's a graph property rather than an instruction. In evaluation scope, where their companion post's turn-level checks run on every conversation for critical behaviors while style preferences get sampled conversation-level scoring. And in the agent's own trade-off instructions, which is what the post means by guiding trade-offs — telling it explicitly that if brevity conflicts with a required disclosure, the disclosure wins.

[Inference] The open question the post leaves is who assigns criticality and whether the scale is graded or binary. Binary is usually enough and avoids a long argument about tiers.

In general, how would you design this system to ensure that callers aren’t able to use the generic LLM capabilities and are limited to its customer support role? What guardrails?

[Inference] Layer them, because no single one holds, and note which layers are constraints versus requests.

The strongest guardrail is architectural, and their design already provides it. If the agent is a set of narrow skills gated by a deterministic state machine (Q27), there is no state in the graph whose job is "answer general questions," so an off-role request has nowhere to go. A monolithic prompt-driven agent has no such structural limit.

The second is capability limitation: the agent can only do what its registered functions do. If no function issues a refund above a threshold, no amount of persuasion produces one. This is the layer that cannot be talked around.

The third is a small, fast input classifier ahead of generation, labelling each incoming turn in-scope, out-of-scope, or adversarial and routing the last two to a fixed response rather than to the model. The fourth is output filtering — the concurrent guardrail model call their latency post describes, running in parallel with the main call so it costs no response time, checking for off-topic content, unauthorized commitments, and data leakage. Compliance disclosures and verification prompts should be templated strings rather than generated text, since there is then nothing to manipulate.

[Inference] The framing that matters: an instruction in a prompt is a request, not a constraint. Anything you genuinely cannot allow should be enforced by what the system is able to do.

How would you improve this system if you were building it today?

[Inference] Five changes, ordered by how much they would move real quality.

First, fix the evaluation loop, because everything else is measured through it. Freeze a holdout set the config is never tuned against, generate fresh scenarios each cycle, and run several seeds per scenario with a significance test rather than eyeballing red versus green. Without this the pass rate is a vanity metric (Q28).

Second, decompose the simulated visitor. Store persona, world state, goal, and expected outcome as four independent artifacts and sample the cross-product instead of hand-building bundled objects (Q17). This is what makes coverage measurable — you can then report which state-machine transitions the scenario pool has never exercised.

Third, build the human-review sampler the post assumes but never describes (Q33): disagreement between two judges, plus uncertainty, plus a small random audit sample to catch the confident-but-wrong turns that uncertainty sampling misses.

Fourth, split automation readiness into two scores (Q10) — build feasibility and automation risk — and let the risk score act as a veto rather than averaging into a single number.

Fifth, add a surgical rollback path. Being able to disable one skill behind a flag, rather than reverting an entire config version, is what turns "rapidly fixes failures" (Q35) from an aspiration into an operational property.

[Inference] The one thing I would keep unchanged is the decentralized state-machine architecture. It is the decision that makes every other part of the system testable.

2025 Blogpost original

Why You Can’t Trust Out-of-the-Box Evaluators

My notes

What I didn’t understand or am unsure about

Question

Answer

“compliant, trusted AI agent”

What are examples of standard compliance requirements for agents?

[INFERENCE] The contact center requirements that actually get encoded as evaluator criteria, grouped by the regime that imposes them:

Debt collection. The FDCPA mini-Miranda ("this is an attempt to collect a debt"), restrictions on contact hours, and no disclosure of the debt to third parties. Regulation F adds call-frequency limits.

Payment handling. PCI DSS forbids storing or logging full card numbers and security codes, which in practice means pausing recording and transcription during card capture.

Health. HIPAA's minimum-necessary rule, plus identity verification before any protected health information is discussed.

Contact and recording. TCPA consent for automated outbound calls, and two-party recording consent in states that require it.

AI-specific. Disclosure that the customer is speaking to a bot, required under California's bot disclosure law and the EU AI Act's transparency provisions, and a working path to reach a human.

The first four predate agents. Only the last is new.

“Enterprises that equate high evaluation scores on OOTB metrics”

What are standard OOTB evaluator metrics?

[INFERENCE] What ships in the common frameworks — RAGAS, DeepEval, Arize Phoenix, LangSmith, and OpenAI Evals — sorted by what each actually measures:

Grounding. Faithfulness, or whether the answer is supported by retrieved context, plus context precision and recall. These are the RAGAS core and the most mature of the group.

Task correctness. Answer similarity against a reference, and task completion.

Safety. Toxicity, PII leakage, bias, and jailbreak or prompt-injection detection.

Conversational quality. Coherence, helpfulness, empathy, tone, verbosity. This is the subjective cluster and the one the post's argument targets.

Agent-specific. Tool-call correctness and trajectory adherence — the newest category and the least standardized.

The post's critique lands hardest on the fourth group, where the construct itself is contested. It applies much less to grounding and tool-call checks, which have something closer to a ground truth.

“Consider the behavior, “Agent Assumes Payment.””

What are other examples of “behavior[s]”?

[INFERENCE] A behavior here is a named, observable conversational act specific enough that a yes-or-no answer is possible. That constraint follows from the post's second principle, and it's what separates a behavior from a quality score. Examples by where they occur:

Opening. Agent verifies identity before discussing the account. Agent states the company and purpose of the call.

Diagnosis. Agent asks a clarifying question before proposing a solution. Agent acknowledges a stated frustration before moving on.

Resolution. Agent commits to a specific next step and timeline. Agent quotes policy without checking the account. Agent over-promises a resolution date.

Compliance. Agent reads the required disclosure before collecting payment details.

Closing. Agent confirms the issue is resolved. Agent offers an escalation path when it isn't.

These read like contact center QA scorecard items, which is almost certainly where the taxonomy comes from — Cresta's origin is human agent quality management.

“ask different stakeholders to label each as a positive or negative example of the behavior”

Are they labeling whether the agent exhibited the behavior or whether it’s good or bad that it did?

[SOURCE] The post is genuinely ambiguous, and you can watch it slide between the two readings. "Positive or negative example of the behavior" is classification language — did this instance instantiate the label. But a few sentences later: "whether it's 'good' or 'bad' depends entirely on the context." That's valence, a different question.

[INFERENCE] The coherent reading is that the labeling task is detection, and the collections-versus-sales contrast is diagnostic of a problem rather than an illustration of the method. In both scenarios the agent performs the same act. Stakeholders resist labeling the sales utterance as "Agent Assumes Payment" because it doesn't feel wrong — their valence judgment is contaminating their detection judgment.

That contamination is worth naming, because it means the disagreement the post attributes to unclear guidelines is partly a task-design flaw. The fix is to split the label into two fields: did the behavior occur, and is it desirable in this context. The first is a classifier; the second is policy. Collapsing them guarantees the inter-annotator disagreement the post then sets out to fix.

“Even among business experts familiar with the space, judgments often differ.”

How to handle disagreements?

[INFERENCE] Four steps, in order:

Measure before resolving. Compute inter-annotator agreement — Cohen's kappa for two annotators, Fleiss' kappa or Krippendorff's alpha for more. The post mentions kappa later for its own evaluators. The number tells you whether the guideline is ambiguous or a particular annotator is an outlier, which have different fixes.

Adjudicate in discussion, not by majority vote. Majority vote produces a label while hiding the ambiguity that caused the split, so the guideline never improves.

Write the resolution down as a rule plus a worked example. The examples transfer better than the abstract rule, and they become the judge's few-shot set later.

Re-annotate a fresh sample to confirm agreement actually rose.

Accept a floor. Some disagreement is irreducible, and there you decide by policy — default to the conservative reading on compliance items — rather than pretending consensus exists.

How to know which/how many examples need multiple reviewers vs single expert?

[SOURCE] Not directly answered, but the post supplies the governing distinction: multiple-annotator alignment is "a key requirement" for models "designed to work across use cases and industries," while aligning to a single expert produces a preference-aligned model that is "easier to define" but has "a more specific scope."

[INFERENCE] That settles the model-level question — cross-customer evaluators need multiple annotators, single-organization evaluators don't. The volume question is separate and the post is silent on it.

You don't multiply-label everything. The normal approach is to double- or triple-label a random subset, often ten to twenty percent, purely to estimate agreement, and single-label the remainder once agreement is acceptable. The subset is a measurement instrument, not the training data.

Beyond that random sample, route specific items to multiple reviewers: those near the decision boundary, those where the first annotator flagged low confidence, and anything in a compliance-critical category where a wrong label is expensive. The logic is stratified sampling — spend extra review where variance is high or the cost of error is high, and accept one pass everywhere else.

I’m not fully convinced that going from a small number of numerical scores to a large number of booleans would solve the inconsistency problem. It feels like we’d have the same amount of room for variation in either case unless we narrow the criteria.

[INFERENCE] Your objection is largely right, and the phrase the post quotes gives it away: the work is being done by precise, not by boolean. Decomposing a ten-point quality score into twenty vague booleans — "was the agent professional" — multiplies the ambiguity rather than removing it. Narrowing each criterion to a concrete checkable fact is the actual mechanism, and that's a change in the construct, not in the response format.

Binary does contribute something real and separable, though. A ten-point scale requires agreement on nine decision boundaries, which the post notes ("one for each decision boundary"). Even with a perfectly clear construct, annotators who agree an example is "good" will still split between seven and eight. Binarizing eliminates that class of noise entirely.

So there are two effects and the post conflates them. Narrowing the criteria is the large one; removing boundary ambiguity is the smaller but genuine one. Neither works without the other.

For Precise Boolean rubrics, what happens if certain criteria aren’t applicable? Like if someone asks an agent for advice, I don’t think grading its factuality would make sense.

[SOURCE] Not addressed. The post presents rubrics as uniformly applied and never mentions inapplicable criteria.

[INFERENCE] You need a third value, and conflating "not applicable" with "no" is a real and common defect. It depresses scores on conversations where the criterion never arose, which makes your aggregate metric a function of conversation mix rather than agent quality — two agents of identical quality score differently because they handled different intents.

Two ways to handle it, which differ in where the logic lives:

Gating. An applicability predicate decides whether the criterion is scored at all. "Was a factual claim made?" gates the factuality check. This is cleaner, because applicability is usually a much easier boolean than the criterion itself.

Three-valued labels. Score yes, no, or N/A, and aggregate as passed over applicable rather than passed over total.

On your advice example specifically, the sharper move is usually to narrow rather than exclude. Advice still rests on checkable premises — "your plan covers this procedure" — so scope factuality to verifiable claims instead of marking the whole conversation inapplicable.

“The only sustainable way to achieve scalable, fully-aligned LLM judges is through continuous human calibration.”

What is the continuous aspect of this?

[INFERENCE] Continuous because three things drift underneath a judge that was calibrated once:

The judge model changes. A provider version upgrade shifts judge behavior, and your agreement numbers were measured against the previous version.

The agent changes. As it improves, the conversations reaching the judge shift toward rarer and harder cases — precisely the region where it was never calibrated.

The standard changes. New products, regulations, and policies redefine what counts as good.

Operationally that means a standing sample of judge outputs re-labeled by humans on a fixed cadence, agreement tracked as a monitored metric with alerting when it falls, mandatory re-validation before any judge model change ships, and newly discovered edge cases folded back into the guideline.

The framing worth carrying: inter-annotator agreement is a production metric, not a one-time acceptance test.

“By injecting human preferences”

What does this mean?

[INFERENCE] Given everything else the post says, this almost certainly means prompt-level calibration: human labels and the clarified written criteria they produce go into the judge's rubric and its few-shot examples. Cresta emphasizes that guidelines are customer-visible and customer-adjustable, which only makes sense if the guideline is text in a prompt rather than weights in a model.

Worth flagging the terminology, because it's borrowed. "Human preferences" in machine learning normally means pairwise comparison data used for RLHF or DPO — you collect A-versus-B judgments, fit a reward model, and optimize against it. Nothing in the post mentions training, reward models, or preference pairs.

The two have very different cost and reproducibility profiles, so the loose usage matters. Rubric calibration is cheap, inspectable, and editable. Preference optimization is none of those. If this comes up with them, it's worth asking which they actually mean.

“Cresta’s guidelines are exposed to customers, so you can understand exactly what each LLM judge is doing, and even adjust the criteria to better fit your organizations’ standards”

In what ways could it make sense to adjust the criteria?

[INFERENCE] Adjustments that legitimately make sense, by what's being changed:

Valence. Keep the detection rule, flip whether the behavior is desirable. This is the collections-versus-sales case exactly.

Strictness. Where the threshold sits — how much hedging still counts as assuming payment.

Applicability. Which conversation types the criterion gates on.

Required elements. Adding a jurisdiction-specific disclosure to a compliant-closing criterion.

Vocabulary. Mapping criteria onto your existing QA scorecard names so results reconcile with the quality program you already run.

One risk the post doesn't raise: adjustable criteria break comparability. Once two customers edit the same evaluator differently, Cresta's "proven to generalize" claim no longer covers either edited version, and cross-customer benchmarks stop meaning anything. Any adjustment should trigger re-validation against human labels, or you've silently decalibrated a judge whose credibility rested on the original guideline.

“OOTB evaluators make big promises about objectivity and scale, but in practice, they deliver metrics detached from meaning.”

This seems plausible but what would be a rigorous quantitative way to measure this?

[INFERENCE] Three tests, distinguished by what each compares against:

Agreement with domain experts. Run the evaluator and your experts on the same conversations and compute kappa or Krippendorff's alpha. Benchmark against human-to-human agreement on that same set, not against one hundred percent — the human ceiling is the only fair target.

Predictive validity. Does the evaluator's score correlate with outcomes that already matter — repeat contact rate, escalation rate, CSAT, compliance audit failures? This is the most direct operationalization of "detached from meaning," and the strongest of the three, because it needs no labels and no one's opinion.

Domain sensitivity. Take matched conversation pairs where experts judge the same behavior differently by context — the collections and sales examples. If the evaluator assigns both the same score, it is demonstrably blind to the thing that determines correctness.

Run all three. The second is where a generic evaluator most often fails without anyone noticing.

2025 Blogpost original

When to Use What: A Practical Guide to AI Agent Testing and Evaluation

My notes

What I didn’t understand or am unsure about

Question

Answer

“where our iteration velocity started to plummet”

What aspect of their testing process was especially slow and would cause this?

[SOURCE] Partly answered. The anti-pattern section names it: they "over-indexed on brittle, turn-specific assertions and fine-tuning before we had enough test volume to see impact × frequency." The trigger was launching complex billing and survey use cases.

[INFERENCE] The mechanism the post leaves implicit is that the expensive step is human, not compute. A turn-specific assertion hard-codes what the agent should say or do at a particular moment, so every prompt or flow change invalidates a batch of them and someone has to rewrite the expectations by hand. Maintenance cost scales with turns per flow times tests per turn, and billing and survey flows are long — which is exactly why those launches were the breaking point.

Two further costs compound it. Brittle assertions produce false failures on valid rewordings, and each one consumes triage time before you learn nothing was wrong. And because the assertions were written before they knew which behaviors were frequent or consequential, much of that maintenance was spent on behaviors that didn't matter.

So the slow part was rewriting and triaging tests between iterations, not running them.

“able to run 15x more tests with a far greater degree of diversity, leading to 35% faster release cycles”

Why does 15x more tests correspond to only a 35% faster release cycle?

[INFERENCE] There's no reason to expect proportionality, because the two numbers measure different things. 15x is test volume — an input. 35% is cycle time — an output with many other components.

The ceiling is structural. If testing occupies half your release cycle, eliminating testing time entirely caps you at 50% faster. A 35% improvement is consistent with testing having been a large but not dominant share of cycle time. Design, implementation, review, and customer acceptance don't speed up because you ran more tests.

More importantly, the 15x isn't what produced the speedup. Running more tests is, on its own, slower. The speedup came from removing the manual maintenance bottleneck in Q1; the extra volume is what that removal made affordable, and it's what bought the quality gain instead.

Worth flagging: this is a vendor figure with no stated baseline or comparison method. Treat the direction as the claim and the magnitudes as unverified.

“improving accuracy by 20%”

What metric is this?

[SOURCE] Never defined. No metric name, baseline, denominator, or measurement window appears anywhere in the post. This is its least substantiated claim.

[INFERENCE] In this domain "accuracy" could plausibly mean any of several things, and they aren't interchangeable: goal or task completion rate (did the customer's issue get resolved), containment rate (handled without human escalation), per-turn correctness against a golden response, or judge-scored flow adherence. A 20% move means something quite different in each.

The ambiguity that matters most is absolute versus relative. Twenty percentage points (70% → 90%) and a 20% relative gain (70% → 84%) are very different results, and the phrasing supports either.

There's also no stated denominator. Accuracy measured over which conversations — all traffic, the complex billing and survey flows that triggered the rework, or the test suite itself? Improvement measured on your own expanded test suite is partly circular, since you changed what you were measuring at the same time.

Without those, the number isn't interpretable. It's worth asking about directly if this comes up in conversation with them.

“Because this can be automated”

Can’t everything they do be automated? How is dynamic testing unique in this regard?

[INFERENCE] Your challenge is correct as stated — static tests also run automatically once they exist. The post is being loose, and the real asymmetry is in a different place: test case generation, not test execution.

Splitting the pipeline cleanly:

Generation. Dynamic testing produces new conversations from a goal or persona specification, so coverage scales with compute. Static testing requires a human to notice a case, capture it, and define expected behavior, so coverage scales with headcount. This is the only stage where the two genuinely differ.

Execution. Both are fully automated. No distinction.

Grading. Both use the same evaluators — deterministic checks and LLM judges. No distinction.

Calibration and review. Human, for both.

So the defensible version of the claim is that dynamic testing is the only method whose coverage grows without proportional human effort, which is why it can serve as the volume driver. The post compressed that into "can be automated," which overstates the contrast.

“from early design exploration through post-launch monitoring though post launch”

What’s the difference between post-launch monitoring and post-launch?

[INFERENCE] So the substantive claim is about cadence, not about two concepts: run dynamic testing across the entire lifecycle, but expect its intensity to drop after launch and spike again around major updates.

The distinction actually worth drawing — which the post does make, later, in its "Bonus Points" section — is between offline testing and real-world monitoring. Offline testing runs simulated or replayed conversations against known expectations. Real-world monitoring evaluates live production conversations where there is no ground truth. The post says the better your production monitoring gets, the less you need high-volume offline testing.

That's the meaningful post-launch split, and it's a different axis from the one the garbled sentence appears to be drawing.

“Running a large number of simulated customer conversations”

How to determine right number?

[INFERENCE] It's a statistical power question, not a fixed number, and it decomposes into two multiplied quantities: runs per scenario, and number of scenarios.

Runs per scenario. You're estimating a pass rate, so precision scales as one over the square root of the run count. Distinguishing a 90% pass rate from a 95% one with confidence takes on the order of several hundred runs per scenario; confirming a flow is roughly working takes tens. Set this from the smallest difference you actually need to detect. Combined with the prior post's pass^k framing, you also need enough reruns to measure reliability rather than capability.

Number of scenarios. Driven by how many distinct intents and flows you support, weighted by the impact × frequency criterion the post uses elsewhere.

A practical stopping rule beats a formula here: keep adding runs until new runs stop surfacing new failure modes and stop moving your pass-rate estimates. That saturation point is observable, and it's where marginal cost exceeds marginal information.

“replay them against updated versions of the agent to confirm nothing regresses”

Is there any reason to expect regressions? What circumstances make it more likely?

[INFERENCE] Yes, and more than in conventional software, because LLM agents lack module boundaries. A prompt is one global object — editing an instruction for the refund flow changes behavior on the billing flow too, since both are conditioned on the whole prompt. There is no compiler or type system to contain the blast radius, and no way to know what you affected except by testing it.

Other sources, roughly in order of how much they matter: model version upgrades, which change behavior everywhere at once; context pressure, where adding instructions degrades adherence to existing ones; tool or schema changes; and retrieval index updates that change what the agent sees.

Circumstances raising the likelihood: large prompt edits, edits to shared or global instructions rather than flow-specific ones, a model swap, and any flow that was only marginally passing before — those flip on small perturbations.

One caution: a single failed replay may be resampling noise rather than a real regression. That's what pass^k is for.

“Use it selectively early on, only for issues caught during dynamic testing or key “golden” flows.”

Why?

[SOURCE] The post supplies the reasoning across two sections. Turn-level assertions are "brittle when used broadly," and the anti-pattern was building them "before we had enough test volume to see impact × frequency."

[INFERENCE] Two distinct reasons, and it's worth separating them.

Shelf life. Early on the agent's design is still moving. A snapshot freezes behavior at a moment in context, so snapshots written during active design encode behavior that's about to be deliberately redesigned. You pay maintenance cost on tests with short useful lives, and worse, they fail for the right reason — you changed something on purpose — which trains the team to ignore failures.

Selection. You don't yet know which behaviors are worth locking down. "Impact × frequency" is the selection criterion, and neither term is observable until you have volume. Dynamic testing is the discovery mechanism that produces that information.

So the sequencing is deliberate: explore broadly with dynamic testing to find out what matters and what breaks, then spend the expensive per-case curation only on what you've learned is consequential and stable.

“Continue building on test cases as a regression set”

How to build?

[SOURCE] Partly answered. The Bonus Points section supplies the intake source: effective real-world monitoring "allows you to discover additional edge cases and errors to add to your existing static test sets."

[INFERENCE] The pipeline that implies: production monitoring flags a suspected failure, a human confirms it's a genuine defect rather than a judge error, the conversation state at the failing turn is captured, expected behavior is defined, and it enters the suite as a snapshot.

Four things the post omits that determine whether the suite stays useful:

Deduplicate on intake, or one recurring bug fills the suite with near-identical cases and skews what you're measuring.

Tag by flow and severity, so you can run compliance-critical subsets on every change and the full suite less often.

Prune deliberately. When behavior is intentionally redesigned, the corresponding snapshots must be retired or updated. Skip this and the suite accumulates false failures until people stop trusting it — the failure mode that kills regression suites.

Record why each case exists, since a snapshot with no rationale can't be safely retired later.

“This can also include adherence to an overall flow.”

I’m assuming this is being evaluated by an LLM. If so, couldn’t we just put the desired flow in the agent instructions as well? In which case, I’d assume it’s very unlikely for it to deviate from that trajectory.

[INFERENCE] The premise — instructed therefore compliant — is where this breaks, for four reasons.

Instruction adherence degrades with load. A flow spec competes with every other instruction in the prompt, and adherence falls as instruction count and context length grow.

Customers force deviation. Side quests and mid-flow topic changes mean the ideal trajectory often can't be walked straight, and the spec underspecifies how to leave and re-enter it.

Instructions conflict. "Be responsive and helpful" pulls against "complete these steps in order," and the resolution isn't deterministic.

Non-determinism. Even a well-followed flow fails some fraction of runs.

There's also a design reason to keep them separate: the instruction states intent, the evaluator measures realized behavior. An agent that can't follow the flow is equally unable to report reliably that it did, so the check has to be external.

“This approach is resilient to prompt changes”

What does this mean?

[INFERENCE] It means the evaluation criterion doesn't depend on the agent's specific wording. A conversation-level check asks whether the refund was issued and the steps happened in the right order. Reword the prompt, and the agent phrases things differently but the answer to that question is unchanged. A turn-level check asserting specific confirmation wording fails the moment the wording changes, even though nothing is actually wrong.

So "resilient" means a low false-failure rate under changes that don't affect correctness — which converts directly into lower test maintenance cost, the bottleneck from Q1.

The trade-off is sensitivity, and the post is right to keep both scopes rather than picking one. A conversation-level check won't catch a missing disclosure at turn four if the outcome still succeeded. Resilience and precision trade against each other here, which is why compliance work stays at turn level despite the maintenance cost.

“takes the right actions at specific steps of an interaction”

How is this different from the static/snapshot testing?

[INFERENCE] So they compose into a two-by-two, and the reason they feel like the same thing is that one cell dominates in practice.

Static plus turn-level is the compliance combination the post recommends — replay a recorded state, assert the disclosure appeared. This pairing is common enough that the two concepts get conflated.

Dynamic plus turn-level is real and useful: run simulated conversations, then check whether the confirmation followed payment capture in each. You get turn-level precision across generated rather than curated inputs.

Static plus conversation-level is also real: replay a recorded conversation and judge whether the whole flow succeeded.

The clean statement: static versus dynamic determines where the customer turns come from. Turn versus conversation determines what you assert on. The first is about inputs, the second about measurement.

“now real world monitoring becomes critical”

What would that consist of?

[INFERENCE] The structural difference the post doesn't name: offline you have ground truth, in production you don't. Everything below follows from that.

Sampling. Judging every conversation with an LLM is cost-prohibitive at production volume. Stratified sampling weighted toward high-risk flows, with full coverage only where checks are cheap and deterministic.

Operational signals that need no judge. Containment rate, escalation-to-human rate, conversation length, repeat-contact rate within 24 hours, and where customers abandon. These are cheap, continuous, and often detect a regression before any judge does.

Drift detection. Alert on movement in those rates rather than on absolute thresholds, since you have no correct answer to compare against.

Feedback loop. Route confirmed failures into the static suite through the Q9 pipeline.

The judge scores tell you about quality; the operational signals tell you fastest that something broke.

2025 Blogpost original

The New World of Non-Deterministic Testing and Evaluation

My notes

Some key insights/new things I learned

  1. Snapshot replay for regression testing can be useful
  2. A broad task an agent is asked to do can be broken down into specific steps whose presence/absence in the agent log can then be used to evaluate its performance
  3. Logic checks like verifying successful API calls can be useful for evaluation

What I didn’t understand or am unsure about

Question

Answer

“Large Language Models (LLMs) introduce non-determinism”

What are the different sources of non-determinism in LLMs?

Two groups.

Deliberate sampling — temperature, top-p, top-k, seed. This you asked for and can switch off.

Non-determinism that survives greedy decoding. Floating-point addition isn't associative, and GPU kernels reduce in whatever order is fastest. That order depends on batch shape, and inference servers batch concurrent requests together, so your output can depend on who else hit the endpoint at the same moment. Mixture-of-experts routing, prefix caching, and kernel version drift add to it. Once any of this flips a near-tied argmax, the divergence compounds through the rest of the generation.

For agents, the model isn't the only stochastic part. Tool responses, retrieval indices, and account state all drift. A perfectly deterministic model would still give non-reproducible agent runs.

“meaning the same input can generate an

infinite number of responses”

Is this true even with nucleus sampling?

False as stated, and nucleus sampling makes it more false, not less.

The output space is finite in every case: a model with vocabulary V and maximum output length L can emit at most V^L sequences. For a 100,000-token vocabulary over a few thousand tokens that is astronomically large, but it is a number. Nothing about sampling makes it infinite.

Nucleus sampling then shrinks it. Top-p truncates to the smallest set of tokens whose cumulative probability reaches p and renormalizes, assigning exactly zero to everything in the tail. So the set of reachable outputs is strictly smaller than under unrestricted sampling. Top-k does the same more bluntly.

Your instinct points the right way — top-p is a counterexample to the framing.

The post's argument doesn't depend on the word, though. It needs only "too many valid outputs to pre-script," which holds under aggressive top-p. The imprecision is rhetorical rather than load-bearing, but worth catching: genuine infinitude would imply things about coverage that a large finite space does not.

“In testing, this can create the illusion of success. For example, an AI agent might tell a customer their account issue is resolved simply because the customer insists it is, even though the agent never confirmed the fix in the system. In a goal-driven simulation, this may appear as a 'pass' even though the task was never actually completed.”

Why would this appear as a pass? Shouldn't the evaluator also have access to the system to check if the issue was resolved?

[INFERENCE] Your objection is right, and the post concedes it two sections later: its own remedy is deterministic evaluation using "logic checks like verifying successful API calls, accurate data retrieval." That is exactly the system access you're describing.

So this isn't inherent to goal-driven evaluation. It's inherent to transcript-only evaluation, and it arises when any of three things hold: the harness has no real backend and tools return canned success; the evaluator is an LLM reading the conversation rather than querying state; or "goal met" has been operationalized as the agent's own closing assertion, making the agent its own judge.

[INFERENCE] The principle: never let the evaluated system supply the evidence for its own evaluation. Assert on backend side effects, not transcript claims. Where a real backend is impractical, give the mock inspectable state so you can still assert on effects.

The post would read better framing this as a design error it has since fixed.

“For example, if an AI agent fails to help a customer reset an alarm, it can be unclear whether it forgot to ask them to reset their WiFi or failed to prompt them to check the battery. And when several breakdowns occur at once, it becomes difficult to understand the full scope of what went wrong. You know it failed, but not where or why.”

In this example, an effective AI agent would have conducted request-specific tasks. Since the set of request-specific tasks would naturally vary for different requests, how can you construct a comprehensive and granular evaluation suite that works for a diverse set of requests?

[INFERENCE] Separate what generalizes from what doesn't, so authoring cost scales with intents rather than conversations.

Generic evaluators, written once. Hallucination, tone, compliance disclosure, whether the agent verified before acting, whether it looped. These apply regardless of what the customer wanted.

Per-intent checkpoint rubrics. For each intent, define the checkpoints correct handling must hit. The alarm example becomes a checklist: check power, check battery, check WiFi, confirm the reset took. Store these as data on the test case, not as evaluator code, so adding an intent is a data change.

Derive rubrics from production. Cluster historical conversations by intent and extract the action sequences characterizing successful resolutions. This is the only part that scales to a long tail — and where Cresta's combined human-agent and AI-agent corpus would actually pay off.

Granularity then falls out: a failure reports which checkpoint was missed.

“Even with sophisticated modeling, simulated testing environments often fail to fully represent real-world customer behavior. LLMs—and even human testers—don't always act like real customers, making it difficult to ensure that test scenarios mirror the conversations agents face in production.”

What does this have to do with goal-driven evaluation specifically?

[INFERENCE] It doesn't belong there, and flagging it is correct. This limits simulation as a testing method, not goal-driven evaluation as an evaluation philosophy. The post itself makes the distinction a paragraph earlier, writing that "simulation became a core testing method" in order to enable goal-driven evaluation — a dependency, not an identity.

The post's own taxonomy separates testing methods (static replay, snapshot, simulated customer) from evaluation approaches (deterministic checks, LLM judges, human review). Goal-driven evaluation is an evaluation approach; simulator realism is a property of one testing method. Run goal-driven evaluation on replayed real transcripts and the realism problem disappears, because the customer turns are real.

[INFERENCE] The defensible version: goal-driven evaluation is most valuable when exploring counterfactual agent behavior, which needs a live counterparty, which in practice means a simulator — so in its highest-value application it inherits simulator bias. Real constraint. Flattening it into "limitation of goal-driven evaluation" hides that the fix lives in the simulator.

“verify behavior remains stable across agent updates”

What behavior is being looked at here and how is it assessed?

[SOURCE] The post specifies the input (real historical or test conversations, held fixed) and the purpose (stability across updates), but names neither the behaviors observed nor the comparison mechanism.

[INFERENCE] Holding the conversation fixed is what makes this tractable: replay the customer turns as a constant and regenerate the agent's side, so any difference is attributable to the agent change rather than to conversational drift.

Observable behaviors, in descending order of how reliably you can assert on them: tool calls (which tool, what arguments, what order), knowledge retrieval (which article or record), policy actions (disclosure read, identity verified), and finally free-text response content.

Assessment has to be mixed, which the post leaves implicit. Structured behaviors admit exact comparison — tool name, argument values, and article ID either match the prior version or they don't. Free text cannot, because exact match would flag every valid rewording as a regression, the precise failure the post's opening section criticizes. So free text needs an LLM judge or a semantic similarity threshold against the prior response.

“Captures a moment in context (e.g., when

the AI agent gave an incorrect or best-practice response) for regression testing”

Why do they call it regression testing? What value does this add compared to the full evaluation suite?

[INFERENCE] The term comes from software engineering, where a regression is a previously-working behavior breaking again. The implied workflow: a bad response is observed in production, fixed through a prompt or flow change, and that exact moment is frozen as a test so the fix can't silently revert when something unrelated changes.

Value over the full suite comes in three parts.

Cost and latency. A snapshot is one turn with fixed context; a simulated conversation is many turns with a second model in the loop. You can run thousands of snapshots per prompt change. You cannot do that with simulations.

Diagnostic precision. A snapshot failure localizes to a known behavior. A simulation failure says the conversation went wrong somewhere — the exact complaint the post raises about diagnostic value.

Institutional memory. The suite accumulates a record of every mistake the system has made.

The limitation is symmetric: snapshots only cover failures you've already seen. Hence the pairing with simulation.

“Uses an LLM-based customer that follows a goal or instruction (e.g., booking a flight) under varied conditions, exposing real-world unpredictability and helping identify how the agent handles edge cases”

How do you actually get the simulation to explore more diverse customer trajectories?

[INFERENCE] The levers, weakest to strongest:

Sampling parameters. Raising temperature on the user model is the obvious move and the weakest — it produces paraphrases, not new trajectories.

Persona and goal sampling. Sample over personas (terse, confused, hostile, changes their mind mid-conversation), goals, and hidden constraints the agent must surface by asking. This is the first real gain, because trajectory diversity comes from intent structure, not word choice.

Seeding from production. Draw opening turns, grievance types, and account states from real transcripts, so you test against the distribution you'll actually see.

Environment perturbation. Inject tool failures, timeouts, missing records, stale data, ambiguous account states. Often the highest-yield source, because it forces recovery paths no user-side variation reaches.

Adversarial or curriculum search. Reward the simulator for finding failures and resample around them. The only approach that systematically finds what you didn't think to look for.

The last two account for most of the real coverage.

“Uses logic checks like verifying successful API calls”

How do you verify successful API calls?

[INFERENCE] This collapses three different assertions, and they catch different failures.

Call level — was the right call made? Assert on the structured tool-call record the agent framework emits: tool name, argument values, ordering, and absence of calls that shouldn't have happened. Assert against the framework's record, not the prose transcript, because an agent can narrate a call it never made.

Response level — did it succeed? Status code, absence of an error payload, and schema validation of the body. Schema validation matters more than it sounds: a 200 with a malformed or empty body is a common silent failure.

Effect level — did the world change? Query the sandbox backend afterward and assert on state: the ticket exists with the right fields, the address is updated. This is the layer that catches the Q3 sycophancy case, because it trusts neither the agent's claim nor the API's response.

Most teams implement the first two. The third needs a sandboxed backend with inspectable, resettable state — real engineering investment, not a framework feature.

“Expert-aligned LLM judge evaluation: Employs calibrated LLM evaluators to assess nuanced behaviors, such as flow adherence, response relevance, or factuality”

Calibrated how exactly?

[INFERENCE] What calibration means in practice:

Build the reference set. Multiple experts independently score the same conversations, then disagreements resolve into a consensus label. Multiplicity matters — one annotator gives you no way to know what agreement is achievable.

Measure agreement. Cohen's kappa or Krippendorff's alpha for categorical judgments; correlation or mean absolute error for scalars. Raw accuracy misleads on skewed label distributions, which compliance checks usually have.

Iterate. Adjust rubric wording, few-shot anchor examples, and thresholds until agreement clears the bar. Anchors do most of the work.

Re-validate on change. Judge behavior shifts with the underlying model version, so this is a standing obligation.

“Manual human review: Involves experts validating edge cases, refining evaluator calibration, and ensuring that LLM-based assessments align with enterprise quality standards”

What is an edge case in this context?

[INFERENCE] Two meanings are in circulation and the post blurs them.

Edge case as rare input — unusual account states, rare intents, unusual customer behavior. The conventional software meaning.

Edge case as contested judgment — ambiguous intent, conflicting policies, conversations where the agent was technically correct but the experience was poor, or cases where an LLM judge and a deterministic check disagree.

The second is what the quoted sentence needs, because it concerns where to spend expert human attention. The operationally useful definition: a conversation where your automated evaluators are unreliable — low judge confidence, evaluator disagreement, or a score near a decision threshold. Those are the cases where human review changes an outcome. Rare-but-unambiguous inputs don't need a human; the evaluator handles them.

This definition also makes human review schedulable rather than open-ended. You route the low-confidence tail to experts, and the volume becomes something you can budget.

How do you determine whether it’s “align[ed] with enterprise quality standards”? What metric do you measure and what threshold do you need to meet?

[INFERENCE] Metric. The same machinery as Q10: agreement between the automated evaluator and expert consensus on a labeled sample. But choose it per evaluator rather than globally, because error costs are asymmetric and differ by type. For a compliance check, recall dominates — a missed disclosure violation is regulatory exposure, a false alarm costs minutes of review. For a subjective quality score, correlation or kappa against consensus. For a factuality judge, precision and recall reported separately rather than blended into an F1 that hides which is failing.

[INFERENCE] Threshold. No universal number, and claiming one would be the wrong interview answer. The principled bar is judge-to-human agreement approaching human-to-human agreement on the same data. If two experts agree 80% of the time, demanding 95% from the judge is incoherent — you'd be asking it to outperform its own ground truth. So measure inter-annotator agreement first and treat it as the ceiling.

For high-stakes evaluators, the usual pattern is a recall floor with accepted lower precision, plus routing everything flagged to human review.

“By combining these layers”

How do you combine turn-based and conversation-based layers? Single score (potentially weighed), multiple scores or what?

[INFERENCE] A single weighted score is the intuitive answer and the wrong one, for two reasons. It lets a strong conversation-level outcome mask a turn-level compliance violation, which is backwards for a regulated contact center. And the weights are unjustifiable — there's no principled exchange rate between "read the disclosure" and "resolved the issue."

What works:

Gating. Hard binary gates on compliance and correctness: any violation fails the run outright, regardless of outcome quality. Graded quality scores are computed only among runs that clear the gates. Some failures are not compensable, and this encodes that.

Scorecard, not scalar. Report per-evaluator pass rates as a vector. Any aggregation happens at the reporting layer and never destroys the components.

Structured roll-up. Turn-level results aggregate through explicit operators — any violation, all required steps present, count of loops — chosen per evaluator. This preserves the drill-down path the post's own diagnostic complaint demands.

“must be assessed for both correctness and quality”

Difference between correctness and quality?

[INFERENCE] Correctness is whether the right thing happened, objectively verifiable: the right API called with the right arguments, the right record retrieved, the issue actually resolved in the system, the compliance step performed. Ground truth you can query. Correctness is measured.

Quality is how it happened, and it requires judgment: clarity, efficiency, tone, whether the customer was made to repeat themselves, whether a real person would come away satisfied. No system state to query. Quality is graded, by LLM judges and human review.

[INFERENCE] The post's own two examples fail in opposite directions and are worth keeping as the canonical illustration:

  • The flight-booking agent caught in a loop of repeated confirmations is correct but low quality — the booking exists, the experience was poor.
  • The sycophantic agent declaring an issue resolved because the customer insisted is superficially high quality but incorrect — it reads smoothly and nothing was fixed.

The second is more dangerous, and it's what a quality-only stack systematically misses. That asymmetry is the argument for keeping both.

“capture both erroneous and "golden" interactions”

Are ‘erroneous’ and ‘golden’ binary or scores?

[INFERENCE] The tension resolves once you see the words doing two different jobs.

As curation labels, binary. Deciding whether a conversation enters the regression suite as known-bad or reference-good is a yes/no promotion decision made by a human. A 0.7 on that axis would mean nothing operationally.

As an evaluation criterion, graded. A Golden Response evaluator compares actual output against the stored reference. Binary exact match would fail on every valid rewording — the precise failure the post opens by criticizing. So it produces a similarity score against a threshold.

So: binary going in, scored evaluating against.

[INFERENCE] A binary curation label also loses information. "Erroneous" flattens a compliance violation and a clunky phrasing into one bucket, which matters if you later want to prioritize the suite or weight failures by cost. A severity field is cheap at curation time and expensive to reconstruct afterward.

My notes

What I didn’t understand or am unsure about

Question

Answer

“But the majority of people who churn out of SNAP do so from failure to successfully navigate the web of actions required to stay enrolled once eligible”

What is this web of actions?

[SEARCH: 7 CFR 273.12 and 273.14] The recurring obligations are four distinct things, each with its own deadline and its own notice:

  1. Change reporting. Under simplified reporting, which most states use, the household must report when gross monthly income crosses the limit for its household size. Under full change reporting, the household must also report unearned income changes over $100, any addition or loss of a household member, residence changes affecting shelter costs, and work hours falling below 20 per week. The deadline is 10 days.
  2. The periodic or interim report. A mid-certification form, typically around month six of a twelve-month period. Failing to return it terminates participation.
  3. Recertification. A notice of expiration, a new application, an interview at least once every 12 months, and verification documents, generally due at least 10 days before the certification period ends.
  4. Work-requirement documentation for able-bodied adults without dependents, which HR 1 extends up to age 65.

[INFERENCE] The word "web" is doing real work here. These are not one process with four steps; they are four independently timed processes whose triggers overlap, and a single life event can put a household into more than one of them at once.

“We conducted a compensated survey with 54 SNAP recipients across 23 states using our voice-based research assistant”

What could the backend of this voice assistant look like?

[INFERENCE] The conventional stack for this is either a pipeline — telephony or in-app WebRTC audio, streaming speech recognition, a language model that holds the interview guide as a loose state machine, then text-to-speech — or a single speech-to-speech model, which cuts the turn latency that makes pipelines feel stilted. The interview guide matters more than the model choice: it has to specify which topics must be covered, when to follow up, and hard guardrails against giving eligibility advice, since a research interview that accidentally advises someone about their case is a liability. After the call you would expect transcription, personally identifiable information redaction, and model-assisted coding of transcripts into themes.

[SEARCH] Several commercial platforms do exactly this kind of AI-moderated interviewing — Listen Labs, Outset, and Strella are the established ones as of 2026. The two metrics the post reports, completion rate and average conversation length, are the metrics such platforms surface natively, though the post does not say whether Propel bought or built.

“Could voice be an effective interface for helping someone report an address change?”

Is the hypothesis that people (a) find it easier to communicate over voice or (b) the address change is more likely to come up over the course of a full conversation as opposed to being proactively volunteered?

[INFERENCE] The evidence in the post supports (b) and points at a third mechanism that I think is the author's actual claim.

Support for (b) is direct: "In several conversations, interviewees reported case changes that had not yet been reported to their state SNAP program." Those disclosures emerged during conversation rather than being proactively volunteered. The strongest single instance is the participant who had given the same address for four years and corrected herself partway through — "I don't live there anymore actually. I'm homeless now." A form would have collected the stale address and recorded it as accurate.

The third mechanism, which the post emphasizes more than either of yours, is translation. Voice lets someone narrate an event in ordinary language and have the system decide which administrative transactions it maps to, instead of requiring the person to pick the correct transaction first. That is about the mapping burden, not about typing ease or recall.

Option (a) is plausible and consistent with the six-minute average completion, but the post offers no evidence for it.

“Although SNAP programs are federally required to deliver notices by mail”

Only by mail or do they do email too?

[SEARCH: FNS guidance on electronic notice options] The claim as written is imprecise. Since November 2017, states may send SNAP notices electronically without a waiver, subject to conditions: the household must consent, the state must provide security protections, and the household can opt back to paper at any time. The permitted combinations are email plus a secure online account, or email plus text. Text alone still requires a waiver, because it gives no delivery confirmation. If an email address becomes inactive, the state must automatically resume paper notices, and paper must remain available on request.

[INFERENCE] The accurate framing is that paper is the default and the mandated fallback, not the only lawful channel. That does not weaken the post's argument so much as relocate it: electronic delivery requires an affirmative opt-in, which is itself one more procedural action that the households most likely to move are least likely to have completed. Brittleness survives the correction.

“Reporting an address change presumes you have an address to report.”

Are there any general mail reception centers that you can put down if you don’t have a specific address? Like a USPS location or something?

[SEARCH] Yes, and a fixed residence is not required for SNAP eligibility. Pennsylvania's SNAP policy manual, which is typical of state manuals, lists as acceptable mailing addresses: a friend or relative's address, a post office box, a social service agency, "any address where the client can get mail without threat of it being lost or stolen," and the county assistance office itself as a last resort.

The USPS option you are describing is General Delivery. Mail addressed to a person's full name, then "GENERAL DELIVERY," then city, state and ZIP, is held at a designated post office for pickup. It is held a maximum of 30 days, pickup requires photo identification matching the name on the mailpiece, and not every post office offers it — availability has to be checked through the USPS location finder.

[INFERENCE] Each of these presumes the person knows the option exists, can reach the pickup point, and holds valid photo ID. Those are the same access constraints that produced the original failure, which is likely why the post frames it as a design problem rather than an information problem.

“we might treat a recent move or other life events as a signal that case continuity is at elevated risk”

Who is ‘we’ here? Propel? If yes, what information do they have access to that would enable this?

[SOURCE] Genuinely ambiguous. The sentence opens "Instead of designing a better address-change process," which reads as a statement to the benefits-design field rather than a Propel product claim. Propel appears elsewhere in the post as the author's employer and as one example of a trusted existing channel, alongside an assistor or a state portal. The post never says what data would drive the risk signal.

[SEARCH] Propel's public product description covers EBT balance on the home screen, transaction tracking, deposit history, and card security features such as locking and blocking out-of-state or online transactions. Propel states it does not store EBT card PINs, complete card numbers, or Social Security numbers, and that it accesses accounts with user permission through state EBT systems. Case status, certification end dates, address of record, and notice history do not appear in that description.

[INFERENCE] So if "we" is Propel, the available signals are indirect: the geography of where transactions occur, a change in deposit amount, or a direct question to the user. That supports a prompt to ask, not a reliable risk score. A state agency, by contrast, already holds the address of record, the certification calendar, returned mail, and notice history, which is where this idea is actually actionable. My reading is that "we" means the field, and the sentence is a design provocation rather than a roadmap.

“we now have an opportunity to use language models that are unusually well suited to translation between highly personalized, unstructured, unique situations (and spoken/written languages) and highly structured administrative systems”

This seems plausible but what is the input layer/UX? Would Propel users have regular AI interviews to get status updates on their life?

[INFERENCE] Recurring scheduled interviews would be a poor design — it would add a new standing obligation, which is the burden the post is trying to remove, and it would collect most of its information when nothing has changed. Three better triggers, which are mutually exclusive and cover the realistic space:

  1. User-initiated. The person opens the app with something to do — "I need to report that I moved" — and the conversation starts from their framing rather than a form.
  2. Signal-initiated. An external cue starts a short check-in: an approaching recertification date, a deposit that changed amount, returned mail, or a long gap in transactions.
  3. Document-initiated. The person photographs or forwards a notice they do not understand, and the assistant explains it and drafts the response.

The post's own requirement that the person confirm and review before submission implies a visible review step, which suggests voice as an input mode inside an app rather than a voice-only product.

“AI could make those systems easier to navigate by absorbing some of that translation work.”

Concrete example?

[INFERENCE, using the reporting rules from Question 1] That one sentence contains three separately reportable items:

  1. A household composition change, since the household she is now part of is not the household on file.
  2. A shelter cost change, which affects the excess shelter deduction and therefore the benefit amount.
  3. An earned income change, which under simplified reporting is reportable only if gross monthly income now exceeds the limit for the household size — and the household size just changed, so the threshold itself moved.

A system absorbing the translation work would recognize all three from the sentence, compute the new income threshold for the new household size, ask only for the hours and wage rate needed to test it, tell her which items actually require reporting and that the deadline is 10 days, name the documents she will need (a lease or rent receipt, recent pay stubs), pre-fill the portal submission, and show her the structured result before anything is sent.

Today, every step of that reasoning is the recipient's unpaid work, and getting it wrong looks identical to not caring.

“It could also make verification, monitoring, and information requests cheaper for institutions, enabling more administrative complexity rather than less.”

Concrete example?

[SEARCH: HR 1 SNAP provisions] HR 1 makes the incentive concrete. Beginning in FY2028, states pay a share of benefit costs scaled to their payment error rate: 5% below a 6% error rate, rising through 15% and 20% to 25% above a 10% error rate. Administrative cost sharing shifts from an even split to 75% state. Work requirements move to ages 17 through 65, the ABAWD exemption age rises from 55 to 65, parents of children aged 7 and above lose their exemption, and exemptions for veterans, people experiencing homelessness, and former foster youth sunset on October 1, 2030.

[INFERENCE] Put those together. A state now has a direct financial reason to drive its error rate down and a much larger population whose work hours must be verified. If an agent can call employers, parse pay stubs, cross-check payroll databases, and generate information requests at near-zero marginal cost, monthly verification becomes affordable where annual verification was the practical ceiling. Every automated request is still a deadline a recipient must meet with documents they may not have. Agency cost falls; the number of procedural failure points rises. One commenter on the post makes the same argument by analogy to automated claims denial in private health insurance — the technology is neutral about which direction it is pointed.

“A person described what happened in ordinary language and the system identified relevant facts”

Does the system require additional context beyond its pretraining/world knowledge? I’m assuming all relevant guidelines are public and digitized, right?

[SEARCH and INFERENCE] Your assumption is mostly right on availability and wrong on sufficiency. Four separate context needs:

  1. Federal rules are public and digitized. 7 CFR Part 273 is on eCFR, and FNS guidance documents are on the USDA guidance portal. Most state policy manuals are also on the open web — Pennsylvania's, which I used for Question 5, is an ordinary public web page.
  2. But state variation is the whole problem. SNAP is state-administered with wide option-level variation: simplified versus full change reporting, certification length, interview waivers, standard utility allowances, ABAWD waiver areas. A model answering from a pre-trained average will be confidently wrong for a specific state, and HR 1 implementation is changing these rules right now, which makes the pre-training cutoff a live hazard rather than a theoretical one.
  3. Some required context is not in any manual. The specific portal's form fields and submission mechanics, what the county office accepts as document intake, and current processing timelines are operational facts, not policy text.
  4. Case-specific state is agency data. The certification end date, what is already on file, which notices were sent and when. No amount of document retrieval provides this; it requires an integration with the state system of record.

So: retrieval over versioned federal and state policy with effective dates attached, plus a read integration with the agency. Pre-training alone produces fluent, plausible, state-inaccurate answers, which is precisely the failure mode a high-stakes benefits system cannot absorb.

2025 Blogpost original

The evaluation engine behind Decagon’s AI agents

How do you evaluate the performance of these agents to see if they’re accomplishing the desired goals?

My notes

Summary

Context

Decagon has AI voice agents for customer support.

Problem

How do you evaluate the performance of these agents to see if they’re accomplishing the desired goals?

Solution

Offline: LLM-as-judge evaluation + Ground truth evaluation

Online: A/B testing

Some key insights/new things I learned

  1. One useful criterion for LLM-as-a-judge evaluation is “Empathy: Does the response demonstrate understanding and care?”

What I didn’t understand or am unsure about

Question

Answer

What are CSAT and resolution rate? How are they measured? Are the numbers available immediately or is there any latency?

From Decagon's own glossary (separate source): CSAT is the share of surveyed customers giving a top rating — their example is 60 of 100 customers rating 4 to 5 out of 5, giving 60% — collected by a short survey delivered "immediately after a meaningful customer touchpoint." Resolution rate is resolved issues divided by total issues times 100, where resolved means the customer needed no further follow-up.

Inferred from background knowledge: the latency question is the important one, and the two metrics behave differently. CSAT arrives within minutes to days, but only from people who respond. Support survey response rates typically sit well below half of contacts, so a per-variant CSAT estimate needs days to weeks of traffic before its confidence interval is narrow enough to act on. Resolution rate is worse, because "no further follow-up" requires waiting out a re-contact window of roughly one to seven days. Until that window closes the measurement is censored, and early readings look systematically better than the settled value. Both are therefore lagging indicators, which is exactly why teams pair them with immediate proxies such as escalation rate, containment, and judge scores computed on live traffic.

“We use models from OpenAI, Anthropic, Gemini, and other providers, alongside our own fine-tuned versions of open-source and commercial models.”

Why so many different models?

Inferred: five reasons, roughly in order of how much they matter.

First, an agent is not one model call. Intent classification, retrieval reranking, guardrail checks, summarization, and final response generation sit at very different points on the cost-latency-quality curve. A small fine-tuned classifier can do intent routing at a small fraction of the cost and latency of a frontier model, and there is no quality reason to use the expensive model there.

Second, enterprise deployments differ in tone, domain vocabulary, and compliance posture, which is what fine-tunes buy you.

Third, their evaluation apparatus makes heterogeneity self-generating. The post says an experiment variant "might involve anything from a new system prompt to an entirely new language model," so if you continuously run model bake-offs per workflow, you end up with a different winner in different places.

Fourth, vendor concentration risk: capacity limits, price changes, deprecations, outages, and data-residency requirements all argue for being able to swap providers.

Fifth, judge independence. Generating with one model family and judging with another avoids the self-preference bias where a model scores its own outputs too highly.

“We rigorously assess every component of the agent, from response generation to retrieval to safety.”

How do you assess safety?

Inferred from standard practice: safety evaluation for a customer support agent breaks into four fairly distinct things.

The first is an adversarial test suite held fixed across releases, covering prompt injection, jailbreaks, requests that would leak another customer's data or PII, out-of-scope advice such as legal, medical, or financial guidance, and brand-unsafe statements. These are scored pass or fail against a written policy, and you measure both directions of refusal: the agent should refuse genuinely unsafe requests and should not refuse benign ones.

The second, and for a support agent the most consequential, is groundedness. The realistic safety failure here is not toxicity but a confident, wrong policy statement, such as telling a customer they qualify for a refund when they do not. You test this by checking that every factual claim in the response is supported by the retrieved context, using either entailment models or an LLM judge given the source material.

The third is action safety. Once the agent can call tools, the question becomes whether it ever took an irreversible action such as issuing a refund or cancelling an order without the required verification.

The fourth is online: guardrail classifiers running inline at inference time, plus continuous sampling of live conversations for review, which is what a Watchtower-style always-on QA product does.

“Every variant must demonstrate consistent performance offline before it is eligible for online testing”

What does consistent mean in this context?

Inferred: two readings are plausible and both are probably intended. The first is consistency across strata, meaning no regression on any sampled slice rather than just a better global average. This reading is supported by the post's emphasis that sampling is stratified across issue types and customer segments; the whole point of stratifying is so that a variant which lifts the aggregate while breaking one customer's billing flow gets caught. The second is consistency across criteria and across repeated runs: no material regression on relevance, correctness, naturalness, or empathy individually, and stable results when the judge is re-run, since LLM judges are stochastic and a single scoring pass has real variance.

In practice this usually gets implemented as a non-inferiority test per slice — the variant must score at least as well as control minus a small tolerance on every gate metric — combined with a genuine win on whatever metric the experiment was designed to move.

“Using an LLM-as-judge system, we evaluate structured triplets consisting of a user query, the context provided to the model, and the model's generated response.”

Only initial user query or intermediate queries as well?

Inferred: almost certainly any turn, not just the first. Two things point that way. The middle element, "the context provided to the model," only has meaning per generation call — it is the assembled prompt for that specific turn, containing retrieved knowledge base passages, conversation history so far, and customer account data. And sampling only opening turns would miss where support conversations actually break down, which is in clarification turns, turns following a tool call, and the decision to escalate.

The limitation worth naming: grading triplets is a static evaluation. Each response is scored conditioned on a history that a human actually produced, so it cannot capture errors that compound across a conversation, where an early turn steers the dialogue somewhere the grader never observes and every later response looks locally reasonable given a history the agent itself would never have reached. The online A/B test is the only part of their described system that catches this, which means failures of that kind are discovered on real customer traffic rather than offline.

“These triplets are drawn from real-world interactions, ensuring that our evaluation process reflects authentic user needs and scenarios”

Sample size and sampling process? How to ensure representativeness/diversity?

Stated in the post: they "sample across key dimensions, such as issue types and customer segments, to ensure balanced testing." No sample size, no selection mechanism, no mention of reweighting.

Inferred: there is a real tension hiding in the word "balanced." Balanced sampling and representative sampling are different objectives. Balancing oversamples rare strata so that each slice has enough data to say something about it, which necessarily distorts the mix away from production. That means the unweighted aggregate score is not an estimate of production quality. To recover a number that describes live performance you have to weight each stratum back by its true traffic share. The post says nothing about doing this, and it is the kind of detail that quietly makes an offline number uninterpretable.

On sizing, the quantity that matters is observations per slice, not the total. As a rough anchor: detecting a 0.1-point shift on a 1-to-5 scale with a standard deviation near 1 needs on the order of 1,600 observations per arm, and resolving a 3-point difference in a pass rate sitting around 90% needs roughly 1,800 per arm. Both numbers drop substantially with a paired design, where the same triplet is scored under both the control and the variant. Since the context is fixed in their setup, pairing is the obvious choice and I would expect them to use it.

“Each response is scored against several key criteria as part of a scalable, continuous, and stratified evaluation process”

Any weighting of criteria? Binary or scores?

Inferred: the criteria are almost certainly scored separately rather than blended, because they are not substitutable. Correctness is a floor and empathy is a preference. If you average them into one number, a warm, fluent, confidently wrong answer can outscore a blunt correct one, which is the single failure an enterprise support vendor cannot ship. The design I would expect, and would argue for, is correctness and relevance as hard gates with pass thresholds, with naturalness and empathy as optimization targets that can only be traded against each other inside the set of responses that already passed.

On the scale: ordinal scores, most likely a 1-to-5 Likert or a three-point pass, borderline, fail scheme, rather than pure binary. Binary throws away information about near-misses, and scales beyond about five points exceed what a judge can apply consistently. Worth knowing that LLM judges are meaningfully more reliable at relative comparisons than absolute ones, so pairwise scoring of variant against control on the same triplet, with position randomized, is the more trustworthy measurement. It does not give you an absolute threshold to gate on, which is probably why an absolute scoring scheme is used here.

If weights do exist, the principled way to set them is to fit them against observed CSAT, so the composite is the thing that actually predicts the business metric rather than a number chosen by intuition.

“We also audit a subset of triplets with human labellers”

What does this mean? What’s being audited?

Inferred: the object under audit is the judge, not the agent. The human label is treated as the reference and the judge's score as a prediction, so the output of the exercise is judge-quality measurement. That means agreement with humans, reported both as raw agreement and as a chance-corrected statistic such as Cohen's or Krippendorff's kappa, and reported per criterion and per slice rather than pooled. Judges are typically strong on relevance and weak on correctness, because correctness requires domain knowledge the judge does not have.

The second function is calibration. If the judge says 92% correct and humans say 85%, you have learned the offset and can either correct the threshold or repair the rubric.

The third is drift detection, and it is the one teams usually discover too late. When the judge model is upgraded, the score scale shifts silently and historical thresholds stop meaning what they meant. A recurring human-audited set is how you notice that the measuring instrument moved rather than the agent.

One thing worth measuring alongside all of this is human-to-human agreement, since it is the ceiling. A judge that agrees with humans as well as humans agree with each other has nothing left to gain.

What model for LLM-as-judge?

Inferred: rather than guess theirs, here are the constraints that determine the choice. It should be a strong model from a family other than the one generating the responses, because judges exhibit self-preference and score their own outputs higher, a well-documented effect in exactly the literature the post cites. It must be pinned to a specific version, because a silent provider-side upgrade moves your metric without the agent changing, and every historical comparison breaks.

Cost is the other constraint. Judging every response with a frontier model is expensive at their volume, so the common pattern is a smaller distilled judge, fine-tuned on frontier-judge plus human labels, doing the bulk scoring, with the frontier model scoring a sample to keep the cheap judge calibrated.

For correctness specifically, a judge without the source material can only assess plausibility, not accuracy. Their triplet format solves this by construction, since the retrieved context is part of what the judge sees.

What prompt/context given to LLM-as-judge?

Inferred: a well-built judge prompt contains six things.

  1. It states the role and presents the three triplet elements under clear delimiters, with an instruction to evaluate only the response.
  2. It gives one rubric per criterion with each score level anchored by a concrete worked example rather than an adjective, because "3 means partially addresses the question" is not reproducible while a demonstrated example is.
  3. It states explicitly that correctness is judged against the provided context, and that a claim unsupported by that context counts as incorrect even when plausible. Without this the judge grades against its own world knowledge and hallucinations pass.
  4. It asks for brief reasoning before the score, then structured output such as JSON so results parse reliably.
  5. It controls for known biases: randomize ordering in any pairwise comparison, never reveal which variant produced a response, and discourage rewarding length, since judges systematically favor longer answers.
  6. It few-shots from the human-audited examples, which has the useful side effect of tying the rubric directly to the human ground truth.

“In parallel, we evaluate responses against a ground truth evaluation set: a curated collection of user queries with ideal responses”

How concretely to evaluate generated responses against ideal responses?

Inferred: the phrase "factuality and intent coverage" describes two different measurements, and they need different machinery.

Intent coverage asks whether the response addressed everything the ideal response addressed. The robust way to measure this is to decompose the reference answer into the discrete facts or steps it contains, then check each one for presence in the generated response. That gives you recall over required elements, paired with a precision check for content the response asserts that the reference does not support.

Factuality is per-claim entailment: extract the claims in the generated response and classify each as supported, contradicted, or unsupported relative to the reference and the source documents.

What not to use is equally important. Exact match, BLEU, and ROUGE punish valid rephrasings and reward parroting, and support answers have many correct phrasings. Raw embedding similarity fails precisely where the stakes are highest, since "you will be refunded" and "you will not be refunded" sit almost on top of each other in embedding space.

One operational issue the post does not raise: ground truth sets decay. When a return window changes from 30 days to 60, the ideal response silently becomes wrong and the benchmark starts penalizing correct answers. The set needs an owner and a refresh cadence tied to policy changes.

“we apply similar rigor across the full spectrum of our AI workflows to ensure each component performs reliably”

How are the other components evaluated?

Inferred, component by component:

  • Retrieval is measured with recall and precision at k and ranking metrics such as nDCG against labeled query-to-document relevance, because generation cannot be correct if the right document was never retrieved. The separate and more useful measurement is context sufficiency: whether the retrieved set contains enough to answer at all. That split matters because it tells you whether a wrong answer is a retrieval failure or a generation failure, and those go to different owners.
  • Intent classification and routing get per-intent precision and recall with a confusion matrix. The expensive errors are misroutes on high-severity intents such as fraud or cancellations.
  • Tool and API calls are scored on recorded traces for whether the right tool was chosen, the arguments were correct, and errors or empty results were handled.
  • Multi-step reasoning and policy adherence are tested with scripted scenarios that have required steps in a required order, such as verify identity, then check eligibility, then act.
  • Escalation is a classifier decision with its own precision and recall. Both error types cost something: unnecessary handoffs cost money, missed handoffs cost trust.
  • End to end, the method that covers what triplets cannot is simulated multi-turn conversations driven by a user-simulator model and scored on task completion.

“We manage rollouts by gradually increasing traffic to the variant group as performance improves”

How concretely to determine when to increase traffic?

Inferred: a defensible ramp has four parts.

There is a fixed ladder of traffic shares, something like 1%, then 5%, 25%, 50%, with a hold at each rung long enough for the lagging metrics to settle, meaning at least one full re-contact window and ideally a full week to average over day-of-week effects.

There are two classes of metric with different decision rules. Guardrails such as escalation rate, safety violations, error rate, and latency are evaluated as non-inferiority tests with automatic rollback when breached, and they move fast enough to read at low traffic. Success metrics such as CSAT and resolution rate are slow and usually only become decisive at higher traffic shares, so they gate the last steps rather than the first.

The statistical point that most teams get wrong: if you look at the data repeatedly and can stop early, ordinary fixed-horizon p-values are invalid and your false positive rate is much higher than you think. The fix is sequential testing — always-valid p-values, mixture sequential probability ratio tests, or group-sequential boundaries — so that continuous monitoring is legitimate rather than decorative.

Finally, promotion should require that no guardrail was breached, not merely that the average improved. An average can absorb a severe regression in one segment, and for a multi-tenant vendor that regression is concentrated on a single customer.

“Customers can also opt out of experimentation, giving them full control over model exposure during evaluation.”

Validity concerns because of selection bias?

Inferred: yes, but it is worth being precise about which kind of validity is threatened, because the two answers are different.

Internal validity survives. Opt-out happens at the customer level and before randomization, and it does not depend on which arm anyone lands in. Randomization within the remaining population is therefore still clean, and the estimate is unbiased for the population that participates.

External validity is what suffers. What you measure is a local effect on opted-in customers, and opt-out is emphatically not random. The customers most likely to opt out are the largest, the most regulated, the most risk-averse, and often the ones with the most complex workflows, which is plausibly the same set where a new model performs worst. So the measured lift is probably optimistic relative to what happens after full rollout, and the bias runs in the direction that flatters the variant.

There is a second problem the post's framing hides. Assignment at the customer level means the effective sample size is the number of customers, not the number of conversations. Conversations within one customer share policies, intent mix, and phrasing, so they are heavily correlated, and conversation-level confidence intervals computed naively will be far too narrow. This needs clustered standard errors or a customer-level analysis.

Mitigations worth asking about: report what share of traffic is opted in, compare the opted-in and opted-out populations on observable characteristics, reweight results toward the full traffic mix, and roll out to opted-out customers in a later, more heavily monitored stage.

“we monitor closely for noise, novelty effects, and performance drift over time”

How are each of these concretely monitored?

Inferred, taking them in order:

Noise is handled by measuring it before you need it. Run A/A tests, splitting control against itself, to learn each metric's natural variance and confirm the randomization is sound, then compute a minimum detectable effect so you do not read differences the design cannot resolve. Variance reduction using pre-period data (the CUPED technique) tightens intervals materially and is nearly free when you have historical metrics per customer. Sequential boundaries keep repeated looks from manufacturing significance.

Novelty effects are detected by plotting the treatment effect against days since exposure rather than pooling the whole period. A novelty or primacy effect appears as an effect that decays or grows over the first days and then flattens, and the rule is to hold the ramp until the curve is flat. Splitting new from returning end users separates the two mechanisms. In a business-to-business agent context it is worth noting that the novelty often sits with the human support agents observing the change rather than with end customers.

Drift is monitored by watching the control arm over time. If control itself moves, something outside the experiment changed: traffic mix, a policy update, an upstream model version, seasonality. Alongside that you monitor input distribution shift in intent mix and query embeddings, and you keep a pinned judge scoring a fixed reference set. If scores on the fixed set move, the measuring instrument drifted rather than the agent, and that distinction is the whole game.

“Insights from every phase feed back into the experiment loop: improving prompts, upgrading models, and refining overall agent behaviors.”

How exactly are these insights used?

Inferred: the version of this loop that actually compounds is error-driven rather than score-driven. A score tells you whether to ship; it never tells you what to change. So the first step is clustering the low-scoring triplets by failure mode — retrieval missed the document, retrieval succeeded but the model answered from priors, tone was wrong, escalation was missed, the tool call was malformed — and sizing each cluster. That converts an uninterpretable three-point score drop into a ranked list of fixes.

Each mode then routes to a different owner. Retrieval misses go to knowledge base content, chunking, or the reranker. Correct-context-but-wrong-answer goes to the prompt or the model. Policy violations go to guardrails. Tone goes to the prompt or a fine-tune.

The mechanism that makes the loop accumulate value is promoting confirmed failures into the ground truth set, so every real defect becomes a permanent regression test and the same bug cannot ship twice. Running the other direction, escalated and low-CSAT conversations from online tests should become the next cycle's offline triplets, which is what keeps the offline set representative as traffic shifts underneath it.

Decagon's own GEPA post is a concrete instance of this loop: prompts optimized against a labeled holdout of roughly 600 examples, reported as a 5 to 6 percent improvement, with the stated discipline being to "treat prompt engineering as software engineering." That is a separate source, not this post.

How would you improve this system if you were building it today?

  1. Add conversation-level evaluation. The largest gap is that offline evaluation is entirely turn-level. Triplet grading conditions each response on a history a human produced, so it structurally cannot see errors that compound across a conversation. Add simulated multi-turn evaluation with a user-simulator model and task-success scoring, so those failures get caught before customer traffic rather than during it.
  2. Treat the judge as an instrument requiring continuous calibration, not a fixed oracle. Report agreement with humans per criterion and per slice, track it over time, pin judge versions, and re-baseline explicitly whenever the judge changes. An uncalibrated judge score is an unvalidated sensor, and the entire offline gate rests on it.
  3. Make correctness a gate rather than a term in an average, for the reason in question 7.
  4. Calibrate offline metrics against online outcomes. This is the measurement I would most want that the post does not mention. The open question is how well an offline score delta predicts the CSAT and resolution-rate delta. You can answer it by backtesting every past experiment, comparing the offline predicted effect against the observed online effect. If that correlation is weak, the offline gate is ceremony. The exercise also yields empirically grounded thresholds instead of chosen ones.
  5. Reweight the stratified sample back to the production traffic mix, so the aggregate offline number estimates something real.
  6. Adopt sequential testing with pre-registered guardrails and automatic rollback, and use clustered inference given tenant-level assignment.
  7. Promote cost and latency to first-class evaluation metrics alongside quality. In live support a three-second response that is marginally worse frequently beats a fifteen-second response that is marginally better, and neither appears anywhere in the post.
  8. Add offline replay of logged conversations against a new variant as a cheap pre-screen before spending any traffic.
  9. Build per-customer dashboards rather than aggregate ones, because a vendor's risk is concentrated per tenant and an aggregate win that hides one customer regressing is a churn risk the average conceals.
2026 Blogpost original

Optimizing GEPA for production: A test-driven approach to prompt engineering

Prompt engineering has traditionally been a manual, iterative process: craft a prompt, test it, refine based on failures, repeat.

My notes

Summary

Context

Decagon has a supervisor model that analyzes conversations and produces structured judgments with reasoning traces. For their customers, the supervisor model is a last line of defense against hallucinated or inconsistent outputs.

Problem

Prompt engineering has traditionally been a manual, iterative process: craft a prompt, test it, refine based on failures, repeat.

Solution

GEPA (Reflective Prompt Evolution) offers a systematic alternative: let an LLM reflect on failures and propose improvements automatically.

Some key insights/new things I learned

  1. GEPA is a systematic approach to improving prompt quality, using an operating model to generate outputs and a reflection model to propose prompt improvements based on shortcomings of the generated outputs
  2. The reflection model for GEPA should probably be a frontier model; lighter-weight models don’t work as well

What I didn’t understand or am unsure about

Question

Answer

“gradient-free optimizer that uses natural language reflection rather than policy gradients to adapt prompts”

What does policy gradient mean?

In RL, the “policy” is the model that chooses outputs. For an LLM, the weights set the probability of every token. A policy gradient method repeats three steps:

  1. Sample outputs from the model for many inputs.
  2. Score each output with a single number, called the reward.
  3. Adjust the weights in the direction that makes high-reward outputs more likely. That direction is the gradient of expected reward with respect to the weights.

GRPO (Group Relative Policy Optimization) is one such method. It samples a group of answers for each input and scores each answer against the group’s average.

GEPA is “gradient-free” because it never changes weights. It edits the prompt text, guided by an LLM’s written diagnosis of what went wrong. A single reward number says how good an output was but not why, which is a big reason policy gradient methods need so many more samples.

“it outperforms reinforcement learning methods like GRPO by up to 20%”

  1. By what metric?
  2. Is it consistently better across most applications or does it vary?

From the GEPA paper (web, arXiv v2, accepted to ICLR 2026):

(a) The metric is each benchmark’s own score on a 0–100 scale, such as exact match, F1 or pass rate depending on the task. The “20%” is a 19-point absolute gap on HotpotQA with Qwen3 8B (62.33 for GEPA versus 43.33 for GRPO), not a relative gain.

(b) It varies a lot. GEPA beat GRPO on five of six tasks, by 19.0, 13.66, 5.19, 2.73 and 0.69 points, and lost on AIME-2025 math by 6 points (32.00 versus 38.00). The average gain was about 6 points. GRPO was run only on Qwen3 8B, with a fixed 24,000 rollouts.

Inference: “Up to 20%” is the best case. Two of the five wins were under 3 points, which may be within noise. The biggest wins came on multi-hop retrieval tasks (HotpotQA and HoVer), where better instructions about what to search for matter most. The one loss came on competition math, where weight training can build reasoning skill that instructions alone cannot.

“35× fewer model rollouts”

What is a model rollout?

Inference (background knowledge): A rollout is one complete run of the system on one input, followed by scoring the result. For Decagon’s supervisor, one rollout means giving the model one conversation with the current prompt, letting it write its reasoning and label, and comparing that label with the ground truth. The term comes from RL, where it means playing out one full episode with the current policy. Rollouts are the natural unit of cost because each one needs at least one LLM call.

From the GEPA paper (web): Budgets are counted in rollouts. GRPO used a fixed 24,000 per task. GEPA’s total budgets ranged from 1,839 to 7,051, and the paper says GEPA reached its best test performance with 4 to 35 times fewer rollouts.

Inference: GEPA’s total budgets were only about 3 to 13 times smaller than 24,000, so the 35× figure must be counted at the point where GEPA hit its best score, not over its full spend. It is a best-case number.

“We applied GEPA to optimize prompts for a production classification task”


What is the actual task? Input, output and objective?

From the post: The model is a supervisor that “analyzes conversations” and produces structured outputs with reasoning traces. It does binary classification, each prediction must be justified by a chain of reasoning, and the data has ground-truth labels (600+ labeled examples). The post calls it “a last line of defense against hallucinated or inconsistent outputs,” but never names the two classes or the metric.

From Decagon’s guardrails post (web, July 2025): Every message Decagon’s AI agent generates is reviewed by a supervisor model before it reaches the customer. The supervisor checks whether the reply stays on topic, reflects the conversation’s intent and sticks to known facts, and it can revise the reply or escalate.

Inference: Assuming this is the same supervisor, the task most likely looks like this:

  • The input is the conversation so far plus the agent’s draft reply, possibly with the knowledge or policy text the agent relied on.
  • The output is a reasoning trace followed by a binary verdict, roughly “OK to send” versus “flag.”
  • The objective is to maximize agreement with human labels on the holdout set, using a prompt short enough to meet the latency target.

“GEPA operates through four main steps”

Is this cyclical or one-time?

From the post: Cyclical. The section is titled “The core loop,” and the post mentions “Each reflection cycle,” “More iterations,” and a reflection model called “~10-20 times during optimization,” which implies roughly 10 to 20 passes.

From the GEPA paper (web): The loop repeats until the rollout budget is used up. Each pass works like this:

  1. Pick a parent prompt from a pool of candidates.
  2. Run it on a small batch of training examples and collect feedback.
  3. Have the reflection model write an improved prompt.
  4. Check whether the new prompt beats its parent on that same batch. Only if it does, score it on the full validation set and add it to the pool.

Parents are sampled from the prompts that achieve the top score on at least one validation example (the “Pareto frontier”), so several different strategies stay alive. At the end, GEPA returns the prompt with the best average validation score.

Inference: The post’s four steps compress a loop that has two gates: a cheap check on the batch, then an expensive check on the full validation set.

“Reflection: A "reflection model" (typically a frontier LLM) analyzes failures and successes, identifying patterns in what works and what doesn't“

All examples fed into a single prompt or multiple?

From the post: Neither all at once nor one at a time. The baseline uses “10 examples per reflection,” and the batch-size ablation tested “how many examples the reflection model should see at once.” So each reflection call sees one batch of 10, and different calls see different batches.

From GEPA’s code, the DSPy docs and the GEPA paper (web): The default reflection prompt packs the current instruction plus, for every example in the batch, its inputs, the model’s response and the feedback into a single call. DSPy’s default batch size is 3, and the paper used 3, so Decagon’s batches were more than three times the standard size. Batches come from the training split, while the validation split is only used for scoring.

Inference: The task model works differently: it is called once per example. Because the reflection model never sees validation examples, validation scores are not contaminated by what the reflector has read.

“Budget: 1.0x multiplier (~150 LLM calls)“

What role does the budget play here? Isn’t the number of calls fixed by the number of examples?

From the post: Budget was one of the seven ablations, with “Diminishing returns beyond 1× baseline.” Calls also scaled with data: about 60 calls for 20 samples and about 150 for 50.

From the GEPA paper (web): The budget caps total rollouts, and the loop keeps cycling until the cap is reached.

Inference: The examples don’t fix the call count, because GEPA is iterative. The examples set the cost of each cycle, and the budget sets how many cycles run. At baseline, the costs are roughly as follows:

  • Scoring the seed prompt on the 50 validation examples costs 50 calls, once.
  • Each cycle costs 21 calls: 10 to run the current prompt on a batch, 1 reflection call, and 10 to score the new prompt on that batch.
  • Each new prompt that wins its batch costs 50 more calls for validation.

So 150 calls buys only about 2 to 4 cycles. That conflicts with the post’s “~10-20” reflection calls, so at least one figure is loose. The “1.0x multiplier” is probably Decagon’s scaling of a budget tied to dataset size, since both 60/20 and 150/50 work out to 3 calls per sample.

“Batch size: 10 examples per reflection”

Does this imply the GEPA process is cyclical/repeated?

From the post: The batch size alone doesn’t prove it, since a one-shot optimizer could also look at a batch. The post confirms repetition elsewhere: “Each reflection cycle sees a batch of examples and proposes refinements,” and the reflection model is called about 10 to 20 times (see #5).

Inference: Yes, and batching is what makes repetition necessary. With 50 training examples and 10 per reflection, no single call sees all the data, so several cycles are needed to learn from all of it. This resembles minibatch gradient descent, which makes many small, noisy updates instead of one update from the full dataset.

Batch size trades off two things. Larger batches let the reflector compare more failures in one call, but they make each reflection prompt longer and more expensive. Smaller batches are cheaper but give a noisier picture of what is going wrong. The post rated this knob “Low impact” and found 10 “sufficient.”

“20–100 is optimal; 500 hurts performance”

Did they test 500 examples with a length-constrained prompt? I wonder if that would fix the overfitting issues.

From the post: No such run is reported. Each experiment “modified exactly one parameter,” and the baseline had no length constraint, so the 500-sample run was unconstrained. It produced prompts 75% longer, performance 2% lower and 10× the compute, which the post attributes to “More iterations with more data.”

Inference: A length cap would likely fix the bloat but only partly fix the performance, for three reasons:

  1. The 500-sample run changed two things at once, more data and more cycles, so we can’t tell which one caused the drop.
  2. Some bloat is built into GEPA. Its default reflection prompt (web, GEPA code) tells the model to “Identify all niche and domain specific factual information about the task and include it in the instruction.” A cap would make the reflector squeeze niche rules into fewer characters rather than stop chasing them.
  3. A 2% drop may be within noise (see #18).

The clean test is a two-by-two design: 50 versus 500 samples, each with and without the 1,500-character cap, holding the number of cycles fixed.

“The optimal range of 20-100 samples provides sufficient diversity”

How are the 20 selected in the first place? Any measures to ensure representativeness?

Inference: It was most likely a random draw, possibly balanced by label. A random 20 carries two distinct risks:

  1. It can miss whole groups. If, say, 10% of conversations should be flagged, a random 20 contains 2 flagged cases on average, and about 1 draw in 8 contains none.
  2. Even with every group present, which specific examples land in the set can swing the result. The +1% gap between 20 and 50 samples could come from this alone.

Better practice would address three separate goals:

  1. For coverage, stratify the draw by label and by known failure type, so every failure mode appears.
  2. For learning signal, over-sample cases the seed prompt gets wrong, since GEPA learns mainly from failures.
  3. For reliability, repeat the run on several different random subsets and report the spread.

“clear inverted-U relationship”

They say U-shape but the numbers below it seem to suggest it’s monotonically decreasing?

From the post: The bullet list starts at 20 samples, the peak, so it only shows the downslope. The table right below it adds a 10-sample row. Relative to the 50-sample baseline, performance was −5% at 10 samples, +1% at 20, the reference at 50, “Similar” at 100 and −2% at 500. That rises and then falls, which is an inverted U. A regular U would fall and then rise.

Inference: “Clear” is generous, for three reasons:

  1. The rising side rests on a single point, the 10-sample run.
  2. The gaps among 20, 50 and 100 samples are 1% or less and likely within noise (see #18), so the top is really a plateau.
  3. Compute grew with sample size, so the curve mixes a data effect with a cycle-count effect.

A more accurate summary is too little signal below 20, a plateau from 20 to 100, and a mild decline by 500. Also, if the frontier models’ “+5–6% improvement” is measured against the seed prompt in the same units, the −5% at 10 samples means that run barely beat the seed prompt.

What metric is being measured by “Performance”?

From the post: It is never defined. The clues are that the task is binary classification with ground-truth labels, that “All configurations were evaluated on a fixed holdout set,” and that results are reported as percentage changes.

Inference: It is most likely accuracy on the holdout set, meaning the share of conversations whose predicted label matches the human label. It could also be F1, which balances precision and recall. Three further ambiguities matter when reading the tables:

  1. The units are unclear. “−2%” could mean 2 percentage points or a 2% relative drop.
  2. The reference points differ by table. The sample-size table compares against the 50-sample optimized prompt, the length table against the unconstrained prompt, and the frontier models’ “+5–6% improvement” most likely against the seed prompt.
  3. The scope is unclear. The model must justify each label, but nothing says whether the reasoning is scored or only the label.

For a supervisor, a missed hallucination is usually the costly error, so recall at a fixed precision would be more informative than accuracy.

Why is compute cost not linear in # of samples?

From the post: According to the table, it is close to linear. Relative to the 50-sample baseline, 20 samples cost 2.5× less and 500 samples cost 10× more, both exactly proportional, and the bullet counts agree (about 60 versus 150 calls). The deviations are small: 10 samples cost 6× less, where 5× would be proportional, and 100 samples cost 1.8× more, where 2× would be. The post is also inconsistent here, since its bullet list says 100 samples cost “2-4× the compute cost.”

Inference: Near-linear cost suggests the budget was set in proportion to dataset size, about 3 calls per sample (see #7). With a fixed budget, cost would have stayed flat. The small deviations most likely come from rounding in the post’s summary, or from how runs end: GEPA works in whole cycles, so a run stops a little under or over its cap, and that slack is a bigger share of a small budget.

“Substantive evolution”

Qualitatively judged or quantitatively?

From the post: Qualitatively, as far as the post shows. The phrase appears only as a table label for GPT-4.1, GPT-5.2 and the Claude models, next to “No change (65 chars)” for GPT-4o-mini. No measure of how much the frontier models changed the prompt is reported.

Inference: The nearby numbers are indirect. The seed prompt was apparently about 65 characters, since GPT-4o-mini left it unchanged at that length. Unconstrained runs reached 5,000+ characters, and frontier reflection models gained +5–6%. So “substantive” likely means the prompt grew and changed a lot, and performance rose. A quantitative version would do three things for each reflection model:

  1. Report how much text changed, such as characters added or edit distance (the number of character edits needed to turn the seed into the final prompt).
  2. Report how much the content changed, such as the number of distinct rules added.
  3. Report what the change achieved, meaning holdout performance versus the seed prompt.

“the reflection model is only called ~10-20 times during optimization, while the task model (the one executing your prompts) is called hundreds of times”

Sure, but the prompt length for the former is much higher, right?

Inference: Per call, the reflection model has the longer prompt. A task call contains the prompt plus one conversation. A reflection call contains the current prompt plus 10 examples, each with its full conversation, the model’s output and feedback, so its input is several times larger, up to about 10×. That means 15 reflection calls can carry as many input tokens as 75 to 150 task calls.

If the task model ran 150 to 300 times at a similar price per token, the reflection model would be roughly a fifth to a half of input cost, not 5–10%. The claim holds only if the task side is much bigger than the post describes, for example many more task calls, long reasoning outputs, or a pricier task model. So the underlying concern is right: call counts understate the reflection model’s share of cost. The post’s call figures also don’t reconcile with each other (see #7).

“the reflection mechanism naturally accumulates details across iterations”

Any concerns about the order of examples affecting prompt?

From GEPA’s code and paper (web): The default reflection prompt contains only the current instruction and the current batch, with no memory of earlier batches, so the prompt text is the only thing carried forward. Two safeguards exist: a new prompt must beat its parent on the batch before being scored on the full validation set, and the candidate pool keeps several lineages alive.

Inference: Yes, order can matter in two distinct ways:

  1. Across cycles, the process is path-dependent. Early batches decide which rules get written first, and later edits tend to add to those rules rather than rewrite them, so a different batch order could produce a different final prompt.
  2. Within a batch, LLMs weigh examples unevenly by position, often favoring the first and last, which can shift which failures the reflector addresses.

The safeguards filter out edits that don’t help, but they don’t remove path dependence. Running “19+” experiments across 7 dimensions works out to about one run per setting, so order effects could explain the 1–2% gaps. Several runs with shuffled order would settle it.

“The challenge is GEPA's default implementation doesn't support length constraints during reflection. We built a custom instruction proposer that encodes the constraint directly into the reflection prompt.”

Is this guaranteed to constrain length as intended?

From the post: The limit lives in the wording of the reflection prompt, and no check on the output is mentioned, even though the overview table calls these “hard character limits.” In practice it seemed to work: the 1,500-character limit produced prompts of about 1,000 characters, and the 500-character limit produced about 400.

Inference: No, not by that mechanism alone. An instruction is a request, not a guarantee. LLMs read tokens rather than characters, so they count characters poorly, and a reflector facing many failures may overshoot. The observed undershoot also means the effective limit was tighter than the stated one. A real guarantee needs a check in code after each proposal:

  1. Measure the new prompt’s length.
  2. If it is over the limit, ask the reflector to shorten it, allowing a few retries.
  3. If it is still over, reject the candidate so it never enters the pool.

From the DSPy docs (web): DSPy’s GEPA exposes an instruction_proposer hook for custom proposers, which is the natural place for this check.

“Better generalization by preventing overfitting to training edge cases”

How do you reconcile this with the -0.8% performance impact?

From the post: All configurations were evaluated on a fixed holdout set “to measure generalization,” so the −0.8% is presumably a holdout result. By the post’s own measure, the cap slightly hurt generalization.

Inference: The claim can be rescued in three ways, but the post directly supports none of them:

  1. It could mean a smaller gap between training and holdout scores, which is less overfitting even with a slightly lower holdout score. Training scores aren’t reported.
  2. It could be a bet that shorter prompts will hold up better on future production traffic, which a fixed holdout can’t show.
  3. The −0.8% could be noise. The pool has 600+ examples and one run used 500, so a holdout kept separate from every run likely has only about 100 examples. At 100 examples, one conversation moves accuracy by 1 point, so −0.8% is less than one conversation. At 85% accuracy, the standard error of a single score (its typical swing from which examples happen to be sampled) is about 3.6 points.

The defensible claim is “the same holdout performance with a 4× shorter prompt,” not “better generalization.”

“achieving 4× compression”


Why is this compression valuable?

Compression pays off in three ways:

  1. It cuts latency. The supervisor sits in the live reply path, and longer inputs take longer to process, so every extra token adds to the customer’s wait.
  2. It cuts cost. Removing about 4,000 characters saves roughly 1,000 input tokens per call (at about 4 characters per token), multiplied by production volume.
  3. It eases upkeep. A person can read and audit a 1,000-character prompt, while 5,000 characters of edge-case rules is hard to review and more likely to contain conflicting rules.

How would you improve this system if you were building it today?

Inference (my recommendations, grouped by stage):

Measurement

  1. Run every configuration several times on a larger, stratified holdout and report confidence intervals, so 1–2% differences can be trusted.
  2. Optimize for the costly error. A missed hallucination usually costs more than a false alarm, so target recall at a fixed precision rather than raw accuracy.

Optimization

  1. Enforce the length cap in code (see #17), and replace GEPA’s default instruction to include “all niche and domain specific factual information” with one that asks for general rules.
  2. Test interactions, such as 500 samples with a length cap at a fixed budget, rather than changing one factor at a time.

Operation

  1. Feed human-reviewed production disagreements back into the labeled pool, and re-optimize on a schedule, promoting a new prompt only if it doesn’t regress on a fixed test set.
  2. At high volume, train a smaller classifier to copy the supervisor’s labels for speed and cost, and route uncertain cases to the prompted model.
2026 Blogpost original

Designing low-latency AI agents through reranker optimization

Latency - when you're speaking with a voice agent, an extra 200–300ms of silence is the difference between a natural conversation and a robotic one.

My notes

Summary

Context

Decagon has voice agents for customer support.

Problem

Latency - when you're speaking with a voice agent, an extra 200–300ms of silence is the difference between a natural conversation and a robotic one.

Solution

During the reranking step, they manipulated the attention mask to score multiple (query, document) pairs in a single batched forward pass

Some key insights/new things I learned

  • Reranking via LLMs can be expedited through batched calls with attention masks
    • (I think) batched calls with attention masks are mathematically equivalent to standalone calls.
  • Distinction between point-wise and list-wise reranking

What I didn’t understand or am unsure about

Question

Answer

“RAG is an often-overlooked aspect of agent architecture that can quietly eat into response time.”

How does the query rewriting stage work for voice conversations? If we want to generate embeddings to retrieve relevant context, how do we structure the lookups?

Inference (background knowledge): In voice, the input is a speech-to-text transcript of the caller’s latest turn. Transcripts are messy, with filler words, misheard product names, and references to earlier turns like “what about the other plan?” A small, fast LLM reads the recent conversation plus that transcript and writes one or a few standalone search queries, resolving the references and fixing likely mishearings. Turns like “yes” or “hold on” can skip retrieval.

Lookups usually follow three steps:

  1. Indexing: split help articles into chunks, embed each chunk, and attach metadata such as customer, product, and language.
  2. Searching: embed each rewritten query with the same model, run nearest-neighbor search filtered by metadata, and run keyword search (BM25) in parallel to catch exact terms like order numbers.
  3. Merging: combine the result lists (for example, with reciprocal rank fusion), remove duplicates, and send the top candidates to the reranker.

A voice-specific speedup is starting retrieval on partial transcripts before the caller finishes.

“reranking typically incurs the largest latency cost, as it must process tens to hundreds of candidate documents”

Are they using an LLM for reranking? If yes, why use that instead of a standard cross-encoder or late interaction Models that are more typical in industry?

From the post: The model is never named. Two details are telling. The post says the latency gains came without replacing the underlying model, and the starting point was a list-wise reranker, which outputs a ranked list. The attention-mask diagram also labels the shared prefix “instruction + query” and draws triangular attention inside each block.

Inference (background knowledge): Very likely yes. Producing a ranked list requires a text-generating model, and triangular (causal) masks are how decoder-only LLMs attend; a BERT-style encoder attends in both directions and would show full squares. The setup matches rerankers like Qwen3-Reranker, which put an instruction and query before each document and score it by the probability of answering “yes.” A point-wise LLM reranker is itself a cross-encoder, since it reads the query and document together. The real contrast is a small bidirectional encoder versus a large causal LLM.

Likely reasons for the LLM:

  1. Quality: LLM rerankers generally beat BERT-size cross-encoders on public benchmarks.
  2. Control: the instruction can carry customer-specific relevance rules.
  3. Continuity: the same model already powered their list-wise setup.

Late-interaction models like ColBERT precompute document token embeddings, so they are faster, but they usually rank less accurately than cross-encoders.

“which drives up latency due to the quadratic time complexity of transformer attention”

Even with local attention/other attention optimizations?

Inference (background knowledge): Mostly yes. Attention optimizations fall into two groups, and neither removes the problem for a typical reranker:

  1. Exact attention with faster kernels, such as FlashAttention, still computes every token-to-token score. Memory grows only linearly with length and speed improves by a constant factor, but compute stays quadratic.
  2. Restricted attention, such as local (sliding-window) or sparse attention, does cut compute to roughly length times window size. But the model has to be trained that way, and list-wise ranking needs long-range comparisons between documents, which a local window undermines.

The post misses two nuances. The quadratic term only dominates once the sequence is longer than about 12 times the model’s hidden size (Kaplan et al., 2020): roughly 12,000 tokens for a hidden size of 1,024, or 49,000 for 4,096. Below that, the feed-forward layers cost more. List-wise reranking also has to generate the ranked list token by token; FIRST (Reddy et al., 2024) reads the ranking from the first output token’s probabilities instead and reports 50% faster inference.

“We recently moved from list-wise to point-wise reranking at Decagon and saw a significant reduction in reranking latency as a result.“

Quality impact?

From the post: No quality results are reported for this switch. The only quality-related claims concern the later packing step: masking keeps point-wise scoring semantics intact and, per the closing section, preserved scoring quality. Neither claim comes with numbers.

Inference (background knowledge): The effect could go either way.

  1. Reasons it could hurt: a point-wise scorer never sees the other candidates, so it can’t judge relevance relative to them or notice near-duplicates. In zero-shot prompting (no reranking-specific training), the Setwise study (Zhuang et al., SIGIR 2024) found point-wise methods very efficient but weak on ranking quality.
  2. Reasons it could be neutral or help: models trained specifically for point-wise scoring, such as Qwen3-Reranker or RankLLaMA, are strong. Point-wise scoring also avoids list-wise failure modes, like sensitivity to the order in which candidates are listed, and it gives each document an absolute score you can threshold.

The way to find out is to measure ranking quality offline (NDCG@10 and recall@k) on labeled support queries, then confirm downstream effects, such as answer accuracy or resolution rate, in an A/B test.

“The overall reranking step is gated on the slowest individual call, and the overhead of many sequential requests adds up quickly.”

What overhead? And why are the requests sequential as opposed to parallelized?

Inference (background knowledge): The sentence mixes two pictures. A step that waits for its slowest call implies requests running in parallel, while overhead that adds up across sequential requests implies requests running one at a time.

The overhead is most likely the fixed cost each request pays regardless of document content:

  1. Transport: serializing the request, the network round trip, and HTTP or gRPC handling.
  2. Scheduling: tokenization, waiting in the server’s queue, and batch formation.
  3. Compute: fixed costs per forward pass (launching GPU kernels, Python dispatch), plus reprocessing the identical instruction and query for every document.

Why sequential: sending requests in parallel from the client doesn’t mean the GPU runs them in parallel. If a model replica can only fit a few sequences per forward pass, the extra requests wait in a queue and effectively run one after another. Rate limits or concurrency caps on the endpoint have the same effect. Adding replicas helps, but it costs GPUs and still leaves the slowest-call problem described in question 10.

“Here's how it works: we concatenate multiple (query, document) pairs into a single packed sequence, separated by a delimiter token. We then modify the attention mask so that tokens within each pair attend only to each other, while attention to tokens from other pairs is masked out entirely. This produces a block-diagonal attention pattern (see diagram below) — each document is scored in isolation, but we amortize the overhead of a model call across many pairs at once.“

Walk me through a brief illustrative example to make this more concrete. Include relevant math.

From the post: The text describes packing (query, document) pairs, but the diagram shows one shared prefix (instruction plus query) followed by delimiter-separated documents, each attending only to the prefix and its own earlier tokens. The example follows the diagram.

Inference (worked example): Take a prefix (q1, q2) and documents A = (a1, a2) and B = (b1, b2), without delimiters. The mask (1 = may attend, 0 = blocked) is:

q1 q2 a1 a2 b1 b2
q1 [ 1 0 0 0 0 0 ]
q2 [ 1 1 0 0 0 0 ]
a1 [ 1 1 1 0 0 0 ]
a2 [ 1 1 1 1 0 0 ]
b1 [ 1 1 0 0 1 0 ]
b2 [ 1 1 0 0 1 1 ]

The only change from a normal causal mask is the four zeros blocking B from A.

Attention is softmax(QKᵀ/√d + M)·V, where M is 0 for allowed pairs and −∞ for blocked ones. Since e^(−∞) = 0, token b1 spreads its weight only over q1, q2, and itself, and A gets exactly zero. So B’s hidden states match a standalone run on q1 q2 b1 b2, as long as b1 and b2 reuse positions 2 and 3 instead of 4 and 5.

Each score is then read at the document’s last token, typically as p(yes) = e^(z_yes) / (e^(z_yes) + e^(z_no)), where the z values are output logits. The post doesn’t specify the readout.

Cost: two separate calls process 8 tokens; the packed call processes 6, because the prefix runs once.

Why do packed scores exactly equal standalone scores?

The claim is that every document token has the same hidden state, at every layer, as the matching token in a standalone run of [prefix, document]. The argument goes layer by layer (induction):

  1. Layer 0. Embeddings depend only on the token and its position. Position encodings are handled separately.
  2. Each next layer. Norms and MLPs act on each token separately, so they keep matching states matched. Attention at a document token mixes only the tokens it's allowed to see. The prefix tokens match because the prefix only ever sees itself. The earlier tokens of the same document match by the previous layer's result. Same inputs through the same weights give the same output.
  3. Final layer. The final logits match, and so does the yes-or-no score.

What would happen if they retained the batch structure but didn’t do the attention masking in the model call?

From the post: This case isn’t tested. The post justifies packing on the grounds that pairs don’t need to interact to get correct scores, and says masking keeps point-wise scoring intact.

Inference (background knowledge): That framing implies removing the mask would break point-wise scoring. With an ordinary causal mask instead, the effects split into quality and latency.

  1. Quality would degrade in two ways. Scores would depend on neighbors: each document’s tokens would attend to every earlier document in its pack, so the same document could score differently depending on what it’s packed with and in what order. At their batch size of 2, the second document in every pack, half of all documents, would be affected. And inputs wouldn’t resemble training data: a point-wise reranker learns from prompts containing one document, so prompts containing several could produce poorly calibrated scores, and later documents would sit farther from the query, adding position bias.
  2. Latency would stay about the same or improve slightly, since the work per forward pass is similar and a standard causal mask can use fast fused kernels like FlashAttention.

In short, you’d keep the speed but lose the guarantee that packed scores equal unpacked scores.

Is attention masking generally available through model provider APIs or would you need to host an open-source model yourself for this to work?

Inference (background knowledge): You would need to self-host, or at least run your own inference code. It depends on the type of API:

  1. General LLM APIs (OpenAI, Anthropic, Google) accept text, not attention masks or position IDs, so this technique isn’t possible through them.
  2. Hosted rerank APIs (for example, Cohere Rerank or Voyage) accept one query plus a list of documents in a single request. The provider can batch internally, which already removes most per-request overhead, but you can’t control the mask or batch size.
  3. Self-hosted open-weight models, on your own GPUs or a dedicated deployment that runs your code, are the only option for custom masks. Hugging Face Transformers accepts custom 4D masks together with position IDs, and PyTorch’s FlexAttention can express shared-prefix and per-document masks while skipping fully masked blocks. Serving engines like vLLM support rerankers such as Qwen3-Reranker but don’t expose a per-request mask, so you’d modify the engine or write your own inference loop.

“With this packing approach in place, the remaining challenge is straightforward: find the optimize batch size that minimizes end-to-end latency for your system.“

Quality?

From the post: Batch size is treated purely as a latency knob, with no quality results by batch size. The post argues packing is valid because pairs don’t interact, and it says masking preserved quality, without numbers.

Inference (background knowledge): In theory, with an exact mask and correct position IDs, each document gets the identical score at any batch size, so quality shouldn’t change. In practice, three things can break that:

  1. Position IDs: if they aren’t reset for each document, later documents sit at different distances from the query, and scores shift with position in the pack.
  2. Numerical drift: different sequence shapes can make GPU kernels add numbers in a different order, which changes logits slightly and can flip near-ties (Thinking Machines, “Defeating Nondeterminism in LLM Inference,” 2025).
  3. Truncation: a maximum packed length can cut off long documents in larger packs.

The check they didn’t report is simple: score the same queries unpacked and packed at each batch size, then compare the largest score difference, rank agreement (Kendall’s tau), and NDCG@10.

“where the overall step is gated on the slowest call”

What does this mean?

From the post: The phrase is part of the case against small packs: with too few pairs per batch you send many requests, and the step waits on the slowest one.

Inference (background knowledge): The reranker sends many requests at once, and the pipeline can’t move on until every one returns. So the step takes as long as the slowest request, not the average one. “Gated” means that one slow call holds the gate shut for everything downstream.

For example, if 20 calls return in 40 milliseconds and a 21st takes 250 milliseconds, the step takes 250 milliseconds.

More calls make a slow one more likely. If each call independently has a 1% chance of being slow, the chance that at least one of m calls is slow is 1 − 0.99^m. That is about 10% for 10 calls and 63% for 100 calls. Dean and Barroso describe this effect in “The Tail at Scale” (2013). Packing more documents into each call reduces m, which is the benefit side of the batch-size trade-off.

“Using too many causes the packed sequence to grow long enough that per-call processing time starts to climb.”

What does this mean?

From the post: This is the other side of the trade-off: more pairs per pack means fewer calls, but longer and slower ones.

Inference (background knowledge): A pack of b documents of L tokens each, after a P-token prefix, is about P + b × L tokens long, and longer inputs take longer to process. “Starts to climb” fits how GPUs behave. At short lengths the GPU has idle parallel capacity, so extra tokens are nearly free; once that capacity is used up, time grows with length, and attention grows with the square of length.

Custom masks make this worse. Unless you use a sparsity-aware kernel such as FlexAttention, attention computes a score for every token pair in the pack and only then discards the masked ones.

Example with P = 100 and L = 300: two documents make a 700-token pack with 490,000 attention scores, while five documents make a 1,600-token pack with 2,560,000. That’s 2.5 times the documents for 5.2 times the attention work, and only about 15% of those scores are allowed by the mask.

Bigger packs are also more likely to contain one very long document, which then becomes the slowest call.

Is there any good theoretical explanation for why a batch size of 2 yields the best latency results?

Inference (background knowledge): Nothing in theory singles out 2; it’s an empirical minimum. A simple cost model does explain the U-shape if total latency tracks total compute (say, calls queuing on a busy GPU). With N documents of L tokens in packs of b:

Latency ≈ N × (F / b + k × L² × b) + constant

F is each call’s fixed cost (overhead plus the shared prefix), and k × L² × b is attention cost that grows as packs lengthen. The first term falls with b and the second rises, so latency is lowest at b* = √(F / (k × L²)). Longer documents push that optimum down.

Batch size 2 beats 1 and 3 when F is 2 to 6 times k × L², meaning each added document is costly relative to a call’s fixed cost. That fits dense custom masks; a sparsity-aware kernel would shrink k and raise the optimum.

With one measurement per point, 1 versus 2 may be noise.

The post seems to focus primarily on the benefits of this approach. Are there any downsides that they haven’t covered?

Inference (background knowledge): The uncovered downsides fall into four groups:

  1. Quality: point-wise scoring can’t see other candidates, so it can’t penalize near-duplicates or reward coverage. Neither change comes with quality metrics, and position-ID handling isn’t described, though getting it wrong silently shifts scores.
  2. Efficiency ceiling: an optimum of 2 still means making half as many calls as there are documents. If dense custom masks are the cause, speed is being left on the table, and the optimum has to be re-tuned whenever the model, hardware, or document lengths change.
  3. Engineering cost: custom masks require self-hosting and custom inference code, can break during library upgrades, and may bypass serving-engine features like continuous batching and prefix caching.
  4. Evidence: without latency values, percentiles, or error bars, the size of the win is unknown. Batch size 5 is slower than the list-wise baseline (per the plot), which goes unexplained, and tail latency (p95 and p99), which matters most for voice, isn’t reported.

How would you improve this system if you were building it today?

Inference (my recommendations): I’d make four changes, roughly in order of payoff:

  1. Send fewer documents to the LLM. Add a cheap first-pass scorer, such as a small cross-encoder or a ColBERT-style model, to cut the tens-to-hundreds of candidates to about 20. Also cache scores for repeated query-document pairs, which are common in support.
  2. Make each pass cheaper. Compute the instruction-plus-query prefix once and reuse its key-value cache (the attention keys and values already computed for those tokens) for every document, and use a sparsity-aware kernel like FlexAttention so cost tracks each document’s own length. That should push the optimal batch size well above 2.
  3. Shape the calls and their timing. Pack by token budget rather than a fixed document count, so no call is much longer than the rest, and start retrieval on partial transcripts while the caller is still speaking.
  4. Measure it. Track p50, p95, and p99 reranking latency, and gate every change on packed-versus-unpacked score equivalence plus NDCG@10 on labeled support queries.

Additional References

  1. A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models
  2. FIRST: Faster Improved Listwise Reranking with Single Token Decoding
  3. Guide: LLM reranking and attention masking from first principles
2026 Blogpost original

How We Rebuilt Playbook Review as a Multi-Agent System

Their original workflow approach had the following issues: Context got lost between LLM calls and subtasks of contract review Edits were scoped by clause Suggestions were made in isolation With these limitations, reviews could…

My notes

Summary

Context

When a company receives a contract from a counterparty, someone in their legal department compares it, clause by clause, against their company’s playbook — an internal rulebook that companies define to codify which terms are acceptable, negotiable, or dealbreakers. Then they mark up the document with redlines, comments, and send it back to the counterparty for review. This goes on until both parties agree to mutually agreeable terms.

Harvey used to have an LLM workflow system to do this review and redlining.

Problem

Their original workflow approach had the following issues:

  • Context got lost between LLM calls and subtasks of contract review
  • Edits were scoped by clause
  • Suggestions were made in isolation

With these limitations, reviews could over-redline, miss flags, insert suggestions in the wrong location, and miss whose paper the contract was on.

Solution

They built a multi-agent system to address these challenges as well as an internal evaluation suite to assess its quality relative to the original system and alternate approaches.

They also built a set of guardrails around the agent team to keep reviews fast, reliable, and within our infrastructure limits.

Some key insights/new things I learned

  • Committee of frontier models for LLM-as-a-judge evaluation

Design Details

  • Architecture
    • Multi-agent (orchestrator + sub-agents); see figure
  • Evaluation
    • Metric(s)
      • Risk classification: Ability to identify and categorize issues by their level of risk.
      • Redline quality: A list of considerations that a revision by a competent lawyer would have, both in substance and style. These rubrics are defined for each example.
    • Data
      • Details unclear but developed through collaboration with in-house legal team
    • Results
      • LLM judges to score against these rubrics. Then used a committee of three frontier models to score independently and aggregate the votes.
      • See “Eval Results“ section in post
  • Deployment
    • Custom timeouts and targeted retries, Limiting concurrency, Prompt caching, Stream results, Using the best model for each task, multi-model harness

What I didn’t understand or am unsure about

Question

Answer

“When a company receives a contract from a counterparty, someone in their legal department compares it, clause by clause, against their company’s playbook”

How long are company playbooks typically? Are they only text or can they contain other kinds of data too?

[Post] Playbooks "can be hundreds of pages long." A single rule has five parts: the standard position (what the company wants), acceptable deviations (what it will settle for), unacceptable deviations (dealbreakers), guidance on how to apply the rule, and whether the clause is allowed to be missing entirely. Their test set covers "hundreds of provisions."

[Search] The real range is wide. Practitioner guides suggest a first playbook is about an afternoon's work covering the ten or twenty terms that account for most of the redlines a team actually sends. A mature playbook for a single contract type is more like thirty to eighty terms across a few dozen pages. Harvey's "hundreds of pages" describes a large company maintaining playbooks for many contract types at once.

[Inferred] They aren't pure prose. A real playbook usually contains a ladder of fallback positions — ask for this, settle for that, walk away at this — along with the exact wording to paste in for each fallback, dollar thresholds that determine which position applies, and a table of who has to approve what. Harvey's five-field structure suggests they store playbooks as structured records rather than as documents to be read. That matters later, because it's what lets them hand one rule to one agent.

“Playbook rules are conceptually simple, and include the following elements”

How generic are these rules? Is there a single playbook for all kinds of contracts?

[Post] There isn't one playbook for everything. The test data "pairs contracts with playbooks for various contract types," so a playbook belongs to a type. Rules are written generically and then interpreted per deal, which is what the guidance field exists for: "how to interpret and apply the rule given specific deal context."

[Inferred] The normal setup is one playbook per contract type — NDAs, master services agreements, data processing agreements, order forms, vendor agreements. They have to be separate because the same term means different things in each. A liability cap that's perfectly sensible in an NDA would be unacceptable in a services agreement worth millions.

Some terms travel well across types: governing law, how notices get delivered, whether the contract can be assigned to an acquirer. Teams often keep a shared set of those and layer type-specific rules on top.

The practical point is that a rule as written is deliberately incomplete. "We prefer mutual indemnification" doesn't tell you what to do on a particular deal. What fills in the blank is the deal context — whose paper the contract is on, which party you represent, how far into the negotiation you are. That's exactly the information the old system kept losing between steps, and that the new one broadcasts to every agent.

“Rules interact across a document”

Is it possible for rules to be contradictory? If so, what happens?

[Post] Yes, and the post is direct about it. In the old system, "two rules could propose contradictory edits to the same sentence, and there was no rewriting for those edge cases" — the user sorted it out by hand. In the new system, when two rules try to change the same text, the merge step flags it and hands it to the lead agent, which writes a single edit "that satisfies all conflicting rules."

[Inferred] Two different situations are being treated as one here, and only the easier one is actually solved.

The common case is two rules that both want to change the same sentence but aren't really opposed. One rule wants a notice period extended from 15 to 30 days; another wants notices delivered by email as well as post. Both can be satisfied by one carefully written sentence, and that's what the lead agent is doing.

The rarer case is two rules that genuinely cannot both hold — one requires uncapped indemnity for data breaches, another caps all liability at the contract value. No wording satisfies both. The post never addresses this. It's really a defect in the playbook rather than in the contract, and the right response is to stop and tell a human, which fits Harvey's own point that "some decisions cannot be delegated." But nothing they describe would detect it, so the lead agent would presumably draft something plausible and move on.

“it's a quality bar that out-of-the-box AI models and agents miss”

This feels fixable by having “The best edit is usually the smallest one” included as instructions in the system prompt. Would that not work?

[Inferred] Adding that line helps, and they almost certainly had something like it already. It doesn't fix the problem, for two reasons.

The first is that telling a model to make a small edit doesn't help it work out which edit is the small one. To change only the three words that matter, you have to know which three words create the legal problem. That's a hard reading task, and an instruction to "be minimal" gives the model no extra help with it.

The second is that the old setup made minimal edits nearly impossible to produce. It asked the model to rewrite the whole clause, then compared the new version to the original afterward and called the difference a redline. Models trained as helpful assistants tidy as they write — they smooth grammar, reorder for clarity, standardize phrasing. Every one of those shows up in that comparison as a proposed change, even though none of them was the point. The instruction to stay minimal was fighting the model's default writing behavior, and losing.

The new system asks for something different: not "rewrite this clause" but "change these specific words in this specific place." That's a change to what the model produces, not a change in how politely you ask for it.

“still produce edits in a timely fashion”

What are the realistic latency constraints/expectations here? Contract review feels like it should prioritize quality over latency.

[Post] 2.6 minutes before, 3.8 after, which Harvey calls "a reasonable tradeoff to make as we are optimizing for the quality of output." A human review takes 20 minutes to five hours. They also added streaming so users "begin reviewing outputs within seconds."

[Inferred] You're right that quality should win, and Harvey clearly agrees — they accepted a 47% slowdown to get the quality gain. But latency still matters, for reasons that have nothing to do with impatience.

The one that matters most is the first ten seconds. If a lawyer uploads a contract and sees a blank screen, they assume it's broken, or they go do something else and lose the thread. Streaming fixes that without making the review any faster overall, which is exactly why they built it.

The second is that nobody runs the review once. A reviewer changes a deal instruction, or decides to push for the standard position instead of the fallback, and runs it again. Five of those at four minutes each is twenty minutes, which puts you back in the range of just doing the review yourself. That's why the system saves each agent's work and reuses it rather than starting from scratch.

So the constraint isn't a hard number. It's roughly: something useful on screen within seconds, and the whole thing done in less time than it takes to get a coffee.

“For each batch of rules, one model call asked: Does the contract match the rule's standard position?”

Are multiple rules being reviewed simultaneously or one at a time?

[Post] Multiple at once, in batches. One model call took a group of rules and asked, for all of them, whether the contract matched the standard position. Rules that didn't match fell through to a second call about acceptable deviations, and those that failed again fell through to a third about unacceptable deviations.

[Inferred] Batching was there to save money. The contract is the expensive part of the prompt — you pay to send the same fifty pages every time — so asking about twenty rules in one call is far cheaper per rule than twenty separate calls.

The cost is exactly what Harvey complains about. All twenty rules share one call's worth of thinking. A rule that requires tracing a definition across four sections gets the same shallow treatment as a rule that a single sentence settles. In their words: "a single LLM call with limited budget won't be able to think through and give accurate output to solve complex documents and playbooks."

The new design flips this. Every rule gets its own agent, and they run at the same time instead of being crammed into one call. The savings that batching used to provide come back a different way: because every agent reads the same contract, the model provider charges full price only for the first one and a fraction of that for the rest. Same economics, without making rules compete for attention.

“Once classifications were done, a separate retrieval step then located supporting text in the document”

How would this work?

[Post] The post doesn't say. We only get that after classification, "a separate retrieval step then located supporting text in the document," feeding a final call that wrote the summary.

[Inferred] The job is to find the exact piece of contract that justifies the classification, so the interface can highlight it and the lawyer can check the reasoning. There are three normal ways to do that.

The simplest is to have the model quote the relevant text, then search the document for that quote to find where it lives. Cheap, and it works right up until the model paraphrases slightly or invents the quote, at which point you find nothing.

The second is semantic search: split the contract into chunks, convert both the chunks and the rule into numerical representations that put similar meanings close together, and return the nearest chunks. This survives paraphrasing but tends to give you a rough region rather than the exact sentence.

The third, and most reliable, is to number every clause up front and have the model return clause numbers instead of text. That's what the new system does everywhere — "we gave each component of the document a unique identifier."

Why it was a separate step at all: the classification call probably didn't have the whole contract in front of it, so the location had to be recovered afterward. And doing it afterward is precisely how you get the failure they describe, where a suggestion lands on the wrong part of the document.

“We diffed the revision against the original, and that diff became the tracked change.”

What does tracked change mean here?

[Inferred] A tracked change is Word's redline — text marked as inserted or deleted, attributed to an author, each one individually acceptable or rejectable by whoever reads it next. That's what the lawyer sees and what goes back to the other side.

What's interesting is that the phrase means two different things in the two systems, and that difference is the whole design.

In the old system a tracked change was an output. The model wrote a new version of the clause, code compared it to the original, and whatever fell out of that comparison became the redline. Nobody could tell which part of the change was the substantive fix and which part was the model tidying grammar along the way.

In the new system a tracked change is what the agent creates. An agent doesn't hand back replacement text; it hands back "in clause 8.2, replace these words with those words, because of rule 14." Once edits are objects like that, you can tell which rule caused which change, detect that two agents want to touch the same words, and later learn from which edits the lawyer kept. None of that is possible when the edit is just a comparison you ran afterward.

“The core dataset pairs contracts with playbooks for various contract types across hundreds of provisions.”

Synthetic or human generated?

[Post] Not stated. Only that they "built our own evaluation suite with our in-house legal team," that the dataset "pairs contracts with playbooks for various contract types across hundreds of provisions," and that the scoring rubrics were written with lawyers.

[Inferred] Almost certainly a mix, with the pieces coming from different places.

The contracts are most likely real ones from public sources. Companies file their material contracts with the SEC, and those filings are the standard source for anyone building legal AI evaluations. Customer contracts usually can't be used, because their confidentiality terms don't permit it.

The playbooks are most likely written by Harvey's own lawyers to pair with those contracts. A customer's real playbook is a competitive asset and generally isn't available for this.

The rubrics are explicitly human — "after sourcing the rubrics with lawyers." The grading itself is done by models.

This matters more than it sounds. If the same organization wrote the playbooks and designed the system, those playbooks will be cleanly structured and unambiguous in exactly the ways the system handles well. Real customer playbooks are messier: inconsistent, partly self-contradictory, written by different people over several years. The improvements are real, but they were measured on the friendlier version of the problem, and the post gives you no way to tell how much friendlier.

“Risk classification: Ability to identify and categorize issues by their level of risk.“

Accuracy, precision, recall?

[Post] Not stated. One number, 59% rising to 77%, and a description of the task as "a classification problem which is a straightforward task."

[Inferred] A single percentage on a task with three possible labels is almost certainly plain accuracy — the fraction of provisions where the system's label matched the lawyer's. Three things make that a weak way to report this.

Most clauses in most contracts are fine. If 80% of provisions match the standard position, a system that labels everything "standard" scores 80% and catches nothing. Accuracy hides that entirely.

The two kinds of mistake also cost wildly different amounts. Flagging something that was actually fine wastes a few minutes of a lawyer's time. Missing a dealbreaker can cost real money. Harvey knows this — their own list of old failures names both "over-redline" and "miss flags" — so they almost certainly track the two separately internally. They just didn't publish it.

And the three labels are ordered: standard, acceptable, unacceptable. Calling an unacceptable clause "standard" is far worse than calling it "acceptable," because in the first case nobody looks at it again. Accuracy scores both as equally wrong.

The number I'd want is how often the system catches the unacceptable ones, reported on its own.

“Redline quality: A list of considerations that a revision by a competent lawyer would have, both in substance and style. These rubrics are defined for each example.“

Binary or score? Weighting of different dimensions?

[Inferred] "A list of considerations… defined for each example" strongly suggests a checklist per example, where each item is a yes-or-no question the grader answers, and the score is the fraction of items passed.

This is the standard way to build these, and it exists to avoid a specific problem. If you ask a model to rate an edit from 1 to 10, you get different answers on different runs and no way to explain the difference. If you ask "does this edit change only the words that needed changing? yes or no," you get something stable enough to compare across versions of the system.

Weighting is then probably implicit rather than declared: a dimension gets more weight by having more checklist items about it. So a rubric with four questions about legal correctness and one about placement effectively weights correctness four to one, without anyone writing down a weight.

The design I'd expect but that Harvey doesn't claim is a hard failure condition on legal correctness. An edit that is beautifully minimal, perfectly placed, and legally wrong should score zero, not two-thirds. Whether they do that changes what 87% actually means, and there's no way to tell from the post.

“We then used a committee of three frontier models to score independently and aggregate the votes.“

Which models?

[Post] Not named.

[Search] Harvey deliberately runs on models from three companies — OpenAI, Anthropic, and Google DeepMind — and their published lineup names Claude Opus 4.6 (for "agentic reasoning and deep multi-step legal analysis"), GPT-5.2, Sonnet 4.6 (for lower latency), and Gemini 3 Flash (for speed and throughput). The obvious reading of "three frontier models" is one top-end model from each of those three companies. (Source: harvey.ai/blog/why-harvey-is-multi-model-by-design)

[Inferred] The part that actually matters isn't which three, it's that they come from three different companies rather than three sizes of the same model.

Models from the same family have been trained on similar data with similar preferences, so they tend to make the same mistakes and like the same writing. Three of them voting doesn't give you three opinions, it gives you one opinion repeated. Spreading across companies gets you genuinely different judges.

That's especially important here because the system being graded is itself built on one of those companies' models, and models rate their own family's output more generously. Mixing providers reduces that, though it doesn't remove it — one of the three judges is still likely from the same family as the system under test. More on this in Q16.

“equipped with tools to search for various key parts of a document”

What kinds of tools? How do they work?

[Post] Three kinds are named: searching the document, looking up rules and other information, and editing the document. The thing that makes all of them work is that "we gave each component of the document a unique identifier," which "allows agents to cite and edit a specific element unambiguously."

[Inferred] The likely tool set breaks into four groups.

Reading and navigating: fetch a specific clause by its number, get the table of contents, read a whole section, follow a cross-reference or jump to where a defined term is defined.

Searching: keyword or meaning-based search over the contract that returns clause numbers rather than raw text, so every result is something the agent can point at precisely.

Editing: propose a change to a specific piece of text in a specific clause, delete something, insert something after a given clause. Each one gets recorded as a redline tagged with the rule that caused it.

Reporting: file the memo at the end with the classification, the position chosen, and the reasoning.

The reason the identifiers matter so much is worth spelling out. Without them, the only way a model can tell you where it means is by quoting the text. That breaks in three routine ways: contracts repeat near-identical language in several places, the model paraphrases slightly when quoting, and sometimes it invents a quote outright. A clause number either exists or it doesn't, and it points at exactly one thing.

“but was very high latency”

What's a reasonable guess for the latency difference between the single-agent and the multi-agent systems?

[Inferred] You can reason it out from the structure. In the multi-agent system, all the per-rule work happens at once, so the total time is roughly: set up, wait for the slowest single rule, merge, final check. It doesn't matter much whether there are 20 rules or 60.

A single agent does the same work one rule after another. With 30 to 60 rules and a couple of minutes of work per rule, you land somewhere between half an hour and two hours — call it ten to thirty times slower. That's not "a bit slow," that's back inside the range of a human doing the review by hand, which is what makes it unacceptable rather than merely annoying.

Two things make it worse than the simple multiplication suggests. The single agent's context grows with every rule it handles, so it gets slower and worse as it goes — the post's own point that quality degrades as you add tokens. And it's one long chain, so a single stuck step stalls everything, with no way to retry just the rule that broke.

Any concerns about using LLM-based systems to both solve and evaluate?

[Post] The risk isn't discussed. What they do describe as safeguards: the rubrics were written by lawyers rather than models, the graders score against fixed rubrics rather than giving free-form opinions, and three models vote independently.

[Inferred] There are real concerns, and they enter at different points.

Models rate their own family's output more generously than other models' output. Using three different providers reduces this but doesn't remove it, since the system under test is presumably running on one of those three families.

More worrying is what all frontier models have in common. They've been trained to prefer writing that is fluent, thorough, and complete. That's precisely the instinct that "lightest touch" is fighting. A grader with that instinct will quietly reward an edit that rewrites the whole clause nicely over one that changes three words, which means the redline score can climb while deals don't actually move any faster.

There's also the plain fact that 53% to 87% was measured against the rubric the team was optimizing toward. Any metric you tune against stops being an independent measurement of the thing you care about.

And the post reports nothing about whether the graders are any good — no figure for how often the model graders agree with lawyers, and none for how often the three models agree with each other. Without those, you can't tell how much of that 34-point gain is real.

What I'd want: lawyers who didn't write the rubrics grading a held-out set blind, plus the rate at which real customers accept the redlines in production.

“lead agent in charge of distributing work between parallel agents”

What does the lead agent tell the parallel agents to do?

[Post] It "spawns a team of subagents to review each playbook rule." Each one receives the shared deal context — "which party they represent, whether the contract is on your template or theirs, how strict or permissive this negotiation should be, and any deal-specific instructions" — plus any precedent documents the user attached. The subagent then "reads its rule, reads or searches the contract, decides whether the contract complies with the playbook, chooses which position to pursue (standard or a specific fallback), and then drafts a list of suggested edits," finishing with a short memo.

[Inferred] So what actually gets handed over is roughly four things: the one rule it owns, in its five structured fields; the shared deal context; a pointer to the contract — its own branch and the clause numbering — rather than the contract text pasted in, which is what makes the caching work; and a description of what to hand back.

What's notably absent is any information about what the other rules found. The subagents run at the same time and are blind to each other. That's a deliberate choice, and it's the reason reconciliation exists at all: the lead agent's job is to divide the work up front and merge it at the end, not to coordinate anyone in between.

“lead agent reconciles the results to ensure final output is of good quality”

What does reconciliation look like concretely? How does it work?

[Post] It happens in two stages. The merge "merges branches the same way a version-control system would: non-conflicting edits apply cleanly to the final doc, and collisions of two rules touching the same text are escalated for the lead agent to resolve." Then the lead agent resolves those conflicts "the way a senior lawyer would by drafting edits that satisfies all conflicting rules," and "takes a final pass to validate the quality of the review as a whole. This ensures that every rule gets classified and has corresponding edits when needed."

[Inferred] The important thing is how little of this involves a model.

Most of it is ordinary code. Each agent's edits are recorded against specific clause numbers and specific character ranges within them. To merge, you check whether any two agents touched overlapping text. If they didn't — which is the case for the large majority of rules — you apply everything and you're done. No model call, no cost, no risk.

Only the overlaps go to the model. For each one, the lead agent sees the original text, the two competing edits, and the memos from both agents explaining what they were trying to achieve, and writes one edit that does both jobs. This is what those memos are for, and it's why they need to be short: they're the only channel through which one agent's reasoning reaches the lead agent.

Then the final pass checks coverage — every rule has a verdict, and every flagged rule has edits — and reads the merged document as a whole.

The gap is that conflicts are found by checking whether two edits overlap in the text. That catches collisions but not everything that can go wrong. See Q20.

“running up to dozens in parallel”

How does it determine how many need to be run? Based on the number of rules?

[Post] Yes — one agent per rule, always: "every rule in a playbook fans out a subagent." The ceiling on how many run at once comes from infrastructure rather than from any planning: "we added concurrency limits and exponential backoffs to avoid overwhelming shared model infrastructure."

[Inferred] So there are two different numbers. The number of agents needed equals the number of rules, which could be in the hundreds — their own test set covers "hundreds of provisions." The number running at any given moment is capped, with the rest waiting in a queue. "Dozens" is describing the cap, not the size of the job.

This is a deliberately simple design, and it's worth noticing what it gives up. There's no step that decides some rules aren't worth an agent, and no grouping of related rules. If a playbook has four separate rules about limitation of liability, four agents each read the same section, each build their own understanding of it, and then three of the four collide with each other and have to be cleaned up afterward.

What they get in exchange is that the behavior is completely predictable, and caching makes all that duplicated reading nearly free. That's a reasonable trade, but it's a trade — I'd revisit it, see Q37.

How does the multi-agent architecture deal with concerns about rules interacting across a document? Especially since they say “[sub-agents] often would need to apply edits to the same paragraph or sentence while addressing different rules”

[Post] Three mechanisms, doing three different jobs.

Agents can't corrupt each other's work, because each one edits its own copy of the document. Without that, one agent's changes would silently overwrite another's — the post's exact concern: "previous work done by one agent would get overwritten by another agent."

Overlapping edits get noticed rather than silently resolved, because the merge step surfaces them as conflicts.

And conflicts get resolved by the lead agent writing a single edit that satisfies both rules, followed by a read of the whole document.

Separately, a single rule that touches four sections is handled inside one agent, since the agent can search the contract and follow references. That's the direct fix for the old system's problem that "edits were scoped by clause."

[Inferred] Sorting out what this does and doesn't cover:

It genuinely solves agents overwriting each other, and it solves an agent being unable to see past the one clause it was handed.

It partly solves two rules wanting to change the same sentence. Whether the combined edit is any good depends entirely on the lead agent, which is working from two short memos rather than everything the two agents actually figured out.

It doesn't solve edits that conflict in meaning without overlapping in text. Suppose one agent narrows the definition of "Confidential Information" in section 1, and a clause in section 11 was relying on the broad version. Those two edits don't touch the same words, so nothing flags them. The same goes for deleting a clause that something else cross-references. Only the final read-through could catch it, and that's one model call over a long document — not a systematic check.

This is the part I'd push hardest on in an interview. The mechanics are genuinely good; the handling of meaning-level consistency is thin.

“we gave each component of the document a unique identifier”

What does this mean concretely?

[Inferred] Concretely: when the contract is loaded, it gets broken into its structural pieces — sections, paragraphs, numbered clauses, list items, table cells, definitions, exhibits — and each piece is given a label that doesn't change, something like sec-8.2 or p-0147.

From then on, everything works in terms of those labels. The document is shown to the model with them attached, so it reads something like "[p-0147] Neither party shall assign this Agreement...". Tools take a label as their argument: read this section, propose a change to that paragraph. Citations are labels, so highlighting in the interface is exact. And checking whether two agents conflict becomes a matter of comparing lists of labels.

The reason this is worth doing is that the alternative — having the model quote text to say where it means — fails in three routine ways. Contracts repeat near-identical language, so "the Notices provision" might match three places. Models paraphrase slightly when quoting. And as edits accumulate, the original text an agent quoted stops matching what's in the document.

One thing that takes care: labels have to stay stable when the document changes. Inserting a new clause between 8.2 and 8.3 needs a fresh label rather than renumbering everything after it, which is why the internal identifiers are kept separate from the clause numbers a reader sees.

Does the orchestrator look at sub-agent output one at a time or all simultaneously in the same prompt?

[Post] Never stated. What we know: each agent finishes by filing "a short memo"; results are streamed as each agent completes; merging is described as working like version control, with only collisions going to the lead agent; and the lead agent "takes a final pass to validate the quality of the review as a whole."

[Inferred] It's almost certainly a mix of three things, and the word "short" in "short memo" is the clue.

Merging isn't a model call at all. It's code comparing which clauses each agent touched.

Conflict resolution happens one conflict at a time. Each one gets its own small model call containing just the original text, the two competing edits, and the two relevant memos. This keeps each call cheap and focused, and lets several conflicts be resolved at once.

The final validation has to see everything, because checking that every rule got handled and that the document reads consistently isn't something you can do in pieces. So the lead agent probably does see all the memos together at that point — and that is exactly why they have to be short. If each agent handed back its full working transcript, sixty of them wouldn't fit in one call.

That's the general shape of this kind of system: the agents doing the reading burn through enormous amounts of text, and only a compressed summary travels back up to the one that has to see the whole picture.

“Every agent in the hierarchy sees this context”

What does “see” mean here?

[Inferred] "Sees" almost certainly means it's written into every agent's prompt, rather than being something an agent has to go and ask for. Two things point that way: it's small and structured, just a handful of fields, and putting shared text at the start of every prompt is precisely what their caching optimization requires.

The likely exception is the precedent documents, which can be large. Those are more plausibly something an agent searches when it needs them than something pasted into sixty prompts.

Why this matters: it's the fix for one of the old system's most embarrassing failures, where reviews would "miss whose paper the contract was on." In a pipeline, that fact could be known at the classification step and simply absent by the time the redline step ran, because each step only got what the previous step passed along. Putting it in front of every agent makes it impossible to lose.

The cost is that you're paying for those tokens on every one of dozens of calls. That's exactly why prefix caching was worth a section of its own.

“persists with the review”

Any concerns about context window dilution?

[Inferred] The concern is right in general, but this design mostly avoids it, for three reasons.

Saved isn't the same as loaded. That state lives in a database attached to the review, not in a conversation that keeps growing. It gets pulled out selectively, which is what "we reuse the relevant agents" is telling you.

Each agent's world stays small. A subagent only ever sees one rule, the contract, and the deal context. If you ask a follow-up about rule 12, rule 12's agent comes back with rule 12's state — not the accumulated state of the other fifty-nine.

And what's saved is summaries, not transcripts. Classifications, chosen positions, edit summaries, reasoning. Not every search the agent ran and every clause it read along the way.

Where it does bite is the lead agent's final pass, which necessarily grows with the number of rules, and a long session where someone keeps iterating on the same review. The post doesn't say whether either of those gets trimmed.

“we reuse the relevant agents”

How is relevance determined?

[Inferred] Relevance is probably worked out structurally rather than by asking a model, and each of the three triggers has an obvious mechanical answer.

Switching the position on a rule is the easy one. You know which rule the user clicked on, so you re-run that rule's agent and nothing else.

A follow-up question sounds harder, but usually isn't, because the question is attached to a specific flagged issue in the interface. The user clicks on the redline and types their question, so you know which rule they mean without having to interpret the text.

The third case is the interesting one. When the user accepts some suggestions, the document changes, and any rule whose edits were aimed at the text that just changed is now working from a stale picture. Since every edit is already recorded against specific clauses, you can compute exactly which rules those are.

That's the second payoff from the identifier scheme — it was introduced to detect conflicts, and it happens to give you invalidation for free. The post doesn't claim they do this, but they've described every ingredient you'd need.

The risk classification scores in the eval results table seem very low. Any secondary review?

[Post] 59% rising to 77%, with no breakdown by category and no mention of any low-confidence re-check. The product is positioned as a first draft — "reliable first pass reviews," "giving users a well reasoned draft as a starting point." There is escalation to a human, but it's for situations where the deal has nuance the playbook can't capture, not for cases where the model is unsure.

[Inferred] Two separate questions here.

Is 77% actually low? You can't tell from the post, and that's the real problem. If it's measuring exact agreement with a lawyer's label on a three-way judgment call, 77% might be close to the ceiling — two competent lawyers won't agree 100% of the time on whether a deviation is acceptable or unacceptable, and the system can't beat a bar nobody has measured. Harvey never publishes how often their own lawyers agree with each other, which leaves the headline number floating without a reference point. That's the single biggest gap in the post.

Is it acceptable for the product? Yes, provided a lawyer reviews everything. A wrong label costs review time, not a bad contract — as long as the mistake is visible. The dangerous case is the clause the system never flagged at all, because nothing surfaces it for anyone to catch. That's the number worth asking about.

On secondary review specifically: the only second look described is the lead agent's final pass, and what that checks is coverage — that every rule got a verdict. Not whether the verdicts were right. An obvious addition would be re-checking the classifications the model was least sure about, or sending anything near the acceptable/unacceptable boundary to a second model. Nothing like that is mentioned.

“The issue would self-resolve usually on retry”

Why?

[Inferred] Because the failure isn't caused by the input, it's caused by the path the agent happened to take.

Model outputs are sampled rather than deterministic, so the same prompt produces a different first move on a different run. These loops almost always start with one bad early step — searching for a term that isn't in the document, misreading which party a clause binds — and then compound from there. Retrying re-rolls that first step.

There's a second effect that reinforces it. Once an agent has made a few failed attempts, those attempts are sitting in its own context, and the model is now predicting what comes next given a history of failure. It tends to produce more of the same. Starting over wipes that history, which is often more important than the re-roll itself.

And some fraction of "stuck" isn't the model at all — it's a slow or degraded response from the provider, which naturally clears.

Note the word "usually." Retrying doesn't rescue a genuinely hard clause, which is why they pair retries with timeouts rather than retrying forever. And retrying one rule without redoing the review is only possible because each agent works on its own copy.

“To mitigate this, we set timeouts by model, review phase, and document size”

What do timeouts in these different categories mean?

[Inferred] These are three inputs to a single time budget, and each one accounts for a different reason a step might legitimately take longer.

Different models work at very different speeds, especially those that do extended internal reasoning before answering. A single threshold would either cut off a slow model doing perfectly good work, or let a fast model hang for minutes before anyone noticed.

Different stages of the review have different natural durations. Classifying one clause against one rule is quick. Reconciling conflicts across dozens of memos is not. Holding both to the same limit means one of them is set wrong.

And bigger documents take longer for the obvious reason — more to read, more sections to search, more steps before the agent is satisfied. A 300-page contract can't be held to a 10-page contract's budget.

So the limit is looked up from those three facts, probably from a table calibrated against how long things actually took during their evaluation runs, and set well above the slowest legitimate case. The point isn't to make anything faster. It's to make sure that when one rule goes wrong, it costs you that rule and not the whole review.

“we added concurrency limits and exponential backoffs”

How to determine parameters of these?

[Inferred] These numbers aren't found by tuning the agent. They come from constraints outside it.

The cap on how many agents run at once is set by the rate limit Harvey has with the model provider, which is shared across everything Harvey does. You work backwards from it: take the tokens per minute you're allowed, divide by the tokens per minute a single agent consumes, and that's the most you can safely have in flight. Then you divide that budget among customers, so one person's 200-rule playbook doesn't consume the entire allowance and stall everyone else's reviews.

Backoff — waiting before retrying after being rate-limited — follows a standard recipe: wait about a second, double it each time, stop doubling somewhere around 30 to 60 seconds, and add a random amount to each wait. That last part is the one people forget and then regret. If forty agents all get rejected at the same instant and all wait exactly two seconds, they all retry at the same instant and get rejected again. Randomizing the wait spreads them out.

The actual tuning is empirical: raise the concurrency cap until you start seeing rejections and total review time stops improving, then settle just below that point.

One interaction the post doesn't address: time an agent spends waiting in the queue shouldn't count against its timeout. If it does, heavy load causes timeouts, which cause retries, which cause more load.

“We used that fact to redesign our prompts to optimize for prefix cache hits”

How to design prompt to enable this?

[Inferred] First, what the caching actually does. When you send a prompt, the provider does expensive setup work on it before generating anything. If the beginning of your prompt is identical to one it processed recently, it reuses that work — charging a fraction of the usual price and skipping the delay. The catch is that it only reuses the part that matches from the very first character, and it stops at the first difference. One changed word near the top wipes out the benefit for everything below it.

That gives you the design rules.

Put the most shared content first and the most varying content last. The order that works is: system instructions, then tool descriptions, then deal context, then the contract, then the specific rule, then the specific task. Any layout that mixes rule-specific text into the contract destroys the shared portion.

Make the shared part character-for-character identical across every agent. No timestamps, no request IDs, no agent number in a header, no data structures that serialize their fields in a different order each time.

And within a single agent's back-and-forth, only ever append to the conversation. Rewriting an earlier turn invalidates everything after it.

The payoff here is unusually large because the contract dominates the prompt and every single agent reads the same one. That's why one change improved both cost and speed.

“we enabled streaming results such that each rule’s result becomes available as soon as its subagent finishes”

Any concerns about quality given lack of orchestrator review at this stage?

[Inferred] Yes, and this is the one place where the post's own argument works against itself. The reason the lead agent exists is that raw output from individual agents can be contradictory. Streaming shows the user exactly that raw output, before anyone has reconciled it.

Three things can go wrong. A user reads rule 7's suggestion, and then reconciliation rewrites it to accommodate rule 23 — so they have to read it again, and may have already made a decision based on the version that's now gone. If they can accept a suggestion before reconciliation finishes, they can accept one half of a pair of edits that were supposed to be resolved together. And watching answers change on screen makes a system feel unreliable even when every change is the system doing its job correctly.

The standard ways to handle this: stream the classifications early, since reconciliation rarely changes whether a clause is a problem, but hold the actual edits until they've been merged. Or show streamed edits marked as still in progress and don't let anyone accept one until the review completes. Harvey's mention of exposing progress suggests something along these lines, but they don't say whether a rule that's about to be reconciled is visibly marked as such.

“Our evaluation showed that no single model performed best across subtasks of contract review”

Does subtasks mean rule classification vs redlining here or is it more granular where different rules get different models?

[Post] By stage of the review, not by rule: "We built the system to switch models across phases of the same review, selecting each one based on its measured quality, latency, and cost for that specific task."

[Inferred] The stages visible in the post are classification, searching and navigating the document, drafting redlines, resolving conflicts, and the final validation pass. These genuinely do want different things. Resolving a conflict between two competing edits is hard reasoning over a small amount of text, which favors the most capable model available. Navigating a 200-page contract to find the assignment provisions is easy reasoning over a lot of text, done many times, which favors whichever model is fastest and cheapest.

[Search] This lines up with how Harvey talks about its model lineup publicly, which sorts models by exactly that kind of profile: Claude Opus 4.6 for "agentic reasoning and deep multi-step legal analysis," Sonnet 4.6 for "strong performance with a significantly lower latency footprint," Gemini 3 Flash for "where speed and throughput matter." That's a by-stage taxonomy, not a by-rule one.

[Inferred] Sending harder rules to stronger models would be a sensible next step, but nothing in the post claims they do it, and it needs something the post never describes: a way to predict which rules are hard before you've reviewed them. More in Q37.

“We built our agent harness to normalize the differences between model providers.”

What does this mean concretely?

[Inferred] It's a layer of their own code sitting between the agent logic and each provider's API, so that the agents are written once against Harvey's own interface rather than three times against three different ones.

The differences it has to absorb are more than cosmetic. The request and response formats differ, including where the system instructions go and what the streaming events look like. Tool calling differs — what the mechanism is called, how tool schemas are written, whether the model can request several tools at once, what happens when it produces malformed arguments. The controls for how much the model thinks before answering differ and don't map cleanly onto each other. Caching differs: one provider wants you to mark explicitly where the cacheable part of the prompt ends, another does it automatically, so a prompt laid out efficiently for one isn't necessarily efficient on the other. And the operational details differ throughout — how rate limits are signalled, what the error types are, how tokens are counted.

This matters more for Harvey than for a typical chat product because their entire scalability story assumes models are interchangeable parts. Picking the best model per stage, setting timeouts per model, backing off per provider — none of that is manageable if each model is its own integration.

The part a harness doesn't solve is that prompts aren't portable. A prompt tuned around one model's habits often does worse on another. So you get the ability to swap models mechanically, but you still have to re-run your evaluations every time you do.

“Every accepted, edited, or rejected redline helps agents understand a legal team's preferences.”

In what way is this feedback actually incorporated?

[Post] It isn't yet. Note the tense: "This will allow us to build towards a personalized system, from users to accounts." No mechanism is described.

[Inferred] There are four ways to do this, and they differ a lot in cost and in how soon you could ship them.

The cheapest and almost certainly first is to use past decisions as examples. Store which redlines the team accepted and which they rejected, and when an agent is about to work on a given rule, pull up how that team handled the same rule before and put it in the prompt. No training required, works per customer, and it uses the same infrastructure as the historical-contract search in Q35.

The second is to improve the playbook rather than the model. If a team rejects your liability fallback every single time, that's not a model problem — the playbook is wrong. Surface it as a proposed change to the rule for a lawyer to approve. This keeps a human in the loop and improves the artifact the customer can actually read and audit.

The third is to train a small model on the accept/reject history and use it to filter out suggestions the team is unlikely to want before they ever appear.

The fourth is fine-tuning the main model on preferences. This is the most powerful and the least likely soon: per-customer fine-tuning doesn't scale, fine-tuning on everyone's data averages away the personalization that was the point, and one customer's contracts can't be used to improve another's reviews.

The complication running through all of these is that a rejection doesn't tell you why. A lawyer might reject a technically correct redline because the counterparty is important and this isn't the fight to pick, or accept a mediocre one because they're behind on the deal. Without capturing the reason, the signal stays noisy — which is probably part of why this is still future work.

“This opens the ability to search across thousands of documents for previous examples to trust agent outputs and quickly understand what a team has approved in the past and why.”

Who is doing this searching and in what part of the pipeline?

[Inferred] The sentence actually describes two different users doing two different things at two different points.

One is the agent, during the review. Before deciding how to handle a provision, it searches the company's past contracts for how that provision was settled before, and lets that shape which position it pushes for. This changes the output, and it's the personalization path.

The other is the lawyer, after the review. That's what "to trust agent outputs" points at: the interface shows a suggestion next to a note saying your team has agreed to this exact language fourteen times before. This doesn't change the suggestion at all, only whether the lawyer believes it.

My guess is they ship the second one first, because the worst case is much milder. A bad retrieval there is a weak citation the lawyer ignores. A bad retrieval in the first case is a bad redline that goes out the door. Both need the same underlying thing: an index of the company's signed contracts broken down by provision, recording what language was finally agreed and whether it was negotiated.

Worth noting what changes between today and that future. Right now the user chooses which precedent documents matter. Once the system chooses, retrieval quality becomes a new way for the review to be wrong.

How would you improve this system if you were building it today?

In rough order of value.

Fix the measurement before touching the architecture. Publish how often two lawyers agree with each other on the same clause, so 77% has something to be compared against. Report how often the system catches the unacceptable deviations specifically, since that's the expensive miss. And have lawyers who didn't write the rubrics grade a held-out set blind. Right now the headline numbers can't be interpreted, and every decision about what to build next rests on them.

Detect conflicts in meaning, not just in text. Their merge catches two edits touching the same words. It can't catch an edit that narrows a definition in section 1 when section 11 depends on the broad version. When the contract is parsed, build a map of which clauses define terms, which clauses use them, and which clauses cross-reference which. Then check the merged document against that map in code. It's a deterministic check, not a model call, and it closes the gap that worries me most.

Group related rules instead of one agent per rule. Four rules about limitation of liability currently produce four agents that each read the same section, each work out the same thing, and then collide. Cluster rules by which part of the contract they target and give one agent the cluster. That removes duplicated work and removes the conflicts before they happen rather than resolving them afterward.

Escalate when the model is uncertain, not only when the deal is unusual. Today escalation is about deal nuance. Add a confidence score to each classification and route the uncertain ones — especially anything sitting near the line between acceptable and unacceptable — to a stronger model or to a person, marked as needing judgment. This is the cheapest way to reduce the risk of a silent miss.

Route by rule difficulty as well as by stage. They pick a model per stage of the review, but difficulty varies far more between rules than between stages. A rough difficulty signal — how long the rule is, how many conditions it has, how many sections it touches — would let the easy majority go to a fast model and concentrate the budget on the hard minority.

Start the feedback loop now, as retrieval. Accept/reject history is the most valuable thing they'll accumulate, and you don't need to train anything to use it. Index it per customer and feed relevant past decisions to the agent as examples. And capture a reason when a lawyer rejects something, even a one-click reason, or the data stays ambiguous forever.

Make streaming honest. Either hold edits back until they've been reconciled, or show them marked as provisional and don't allow accepting one until the review finishes. Otherwise streaming quietly gives away the consistency guarantee the lead agent exists to provide.

2025 Blogpost original

Effective context engineering for AI agents

Large contexts can inhibit LLM performance.

My notes

Summary

Problem

Large contexts can inhibit LLM performance.

Solution

  • Compaction
  • Structured note-taking
  • Sub-agent architectures

Some key insights/new things I learned

  • Structured note-taking, or agentic memory, is a technique where the agent regularly writes notes persisted to memory outside of the context window. These notes get pulled back into the context window at later times.
  • Compaction in Claude Code retains five most recently accessed files in context
  • Compaction in Claude Code clears tool calls and results

What I didn’t understand or am unsure about

Question

Answer

Is the main difference between prompts and context that context is information + instructions and prompts are just instructions?

Yes, pretty much. Context is also often dynamic.

“we need strategies for managing the entire context state (system instructions, tools, Model Context Protocol (MCP), external data, message history, etc).“

In what form is MCP included in the context?

MCP isn't a protocol object in the window — it gets flattened into ordinary text. Three forms, MECE by when they land:

  1. Connection time — tool schemas. Each server advertises tools; the client renders each as name + description + JSON Schema, injected into the tool-definitions block. This is the dominant cost: tens to hundreds of tokens per tool, paid every turn. Twenty servers × 15 tools is easily 20k+ tokens before you say anything.
  2. Connection time — server instructions, resources, prompt templates. Servers can return an instructions string (typically appended near the system prompt) plus listings of available resources and prompt templates.
  3. Call time — tool results. Whatever content blocks the server returns land in the message history and persist there.

Point 1 is why MCP makes the post's "bloated tool sets" failure mode the default rather than a mistake you have to commit.

“This results in n² pairwise relationships for n tokens.”

Is this true for modern architectures as well? What about local attention?

True of the mathematical object, increasingly false of the implementation. Three things to keep apart:

  1. Compute vs. memory. FlashAttention removed the n² memory cost (it never materializes the attention matrix) but keeps n² FLOPs. GQA and MLA shrink the KV cache, not the pairwise count.
  2. Sparse and local attention genuinely break n². Sliding-window attention (Mistral, Gemma) costs n·w. Frontier models interleave — a few global layers among many local ones. DeepSeek's NSA learns which blocks to attend to. [Search] 2026 work is predominantly hybrid: sparse + linear + a few full-attention layers.
  3. It doesn't rescue the conclusion — it strengthens it. A token in a sliding-window layer cannot see distant tokens directly; it depends on multi-hop propagation up the layer stack. So long-range degradation is, if anything, more likely under local attention. The post's "attention budget" claim survives even though its stated mechanism is a simplification of what frontier models actually run.

“Additionally, models develop their attention patterns from training data distributions where shorter sequences are typically more common than longer ones.”

Don’t models typically get fine-tuned on representative data lengths even if pre-training data has more shorter sequences?

[Post] Doesn't address it — just asserts the pretraining-distribution claim.

[Inferred] Your objection is correct and the post's framing is dated. The standard modern recipe is:

  1. Pretrain short (4k–8k) for most of the token budget — cheap.
  2. A long-context extension phase: continued pretraining at 32k → 128k → 1M, using RoPE base frequency scaling (ABF) or YaRN, on data deliberately upsampled for long documents and synthetic long-range tasks.
  3. Long-context SFT / RL on tasks requiring long-range dependence.

Three things blunt the correction:

  1. Volume. The long-context phase is typically low single-digit percent of total training tokens. Parameters are still overwhelmingly shaped by short-sequence gradients.
  2. Data quality. Genuine documents with real long-range dependencies are scarce. Much long-context training signal is constructed — concatenated documents, synthetic needle tasks — which teaches retrieval more than reasoning across the whole span.
  3. Measurement. Needle-in-a-haystack is near-saturated, but multi-hop and aggregation-over-long-context benchmarks (RULER and its successors) still degrade sharply with length.

Net: the post's mechanism is stated too simply; the empirical claim holds for reasons that survive your correction.

What’s a two sentence summary of position encoding interpolation?

Rather than asking the model to handle position values it has never seen, you rescale the positions of a longer sequence down into the range it was trained on — position 8,000 in a 32k input is fed as if it were position 2,000 in the original 8k range. This works because interpolating inside the trained range is far safer than extrapolating outside it, but it compresses the positional signal so adjacent tokens become harder to distinguish, which is the "degradation in token position understanding" the post mentions.

“smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome”

Dual objective optimization or satisficing?

[Post] Written as if it were a joint optimum, and never resolved. The post does give itself an escape hatch elsewhere: "minimal does not necessarily mean short."

[Inferred] As literally written it's ill-posed — you cannot minimize tokens and maximize success without a stated exchange rate. Three readings:

  1. Scalarized dual objective. max P(success) − λ·tokens. Coherent, but λ is never stated, and in practice λ is tiny next to the cost of a task failure.
  2. Constrained. min tokens subject to P(success) ≥ threshold. This is what "smallest possible set" implies grammatically, and it is the defensible reading.
  3. Satisficing, in practice. Nobody maps the frontier. Teams reach "good enough," trim context until quality visibly drops, and stop.

[Inferred] The more useful reframe: it is effectively lexicographic — quality first, tokens as tiebreaker — because over the relevant range the two objectives are not opposed. Adding low-signal tokens hurts both. A real tradeoff only appears once every remaining token is genuinely informative, which most production prompts never reach.

“System prompts should be extremely clear and use simple, direct language that presents ideas at the right altitude for the agent.”

How good are LLMs at writing system prompts for other LLMs? Could I give an LLM examples of good and bad system prompts and expect it to come up with a good system prompt for my system?

[Inferred] Split into two capabilities, because the answer differs sharply:

  1. From examples of good/bad prompts alone — weak. Style transfer works; you'll get something that looks like a good prompt. But prompt quality isn't a property of the text, it's a property of text × task × model. With no task data the model has no signal about your failure modes. Expect a plausible first draft, not a good prompt.
  2. From examples of good/bad task outputs plus a metric — strong. This is the real answer, and it is a solved-enough tooling problem. [Search] The two current standards are DSPy's MIPROv2 (bootstraps demonstrations, proposes candidate instructions, Bayesian search over combinations) and GEPA (reflective evolutionary search — the model reads execution traces and failures, then mutates the prompt). GEPA generally wins at lower rollout budgets. TextGrad and OPRO are the earlier academic line.

[Inferred] Concrete path, matching the post's own recipe: hand-write a minimal prompt → build a 30–100 case eval set with a scorer → run GEPA or MIPROv2 against it. The eval set is the work; the optimizer is cheap. Without evals you are doing vibes whether a human or an LLM writes it.

“then add clear instructions and examples to improve performance based on failure modes found during initial testing”

How to address risk of overfitting?

[Post] Never mentions the risk. This is a real gap — as written, the recipe is textbook dev-set overfitting.

[Inferred] Mitigations, MECE by where you intervene:

  1. Data. Hold out a test set you never inspect during iteration. Refresh the dev set from production traffic periodically. Keep a regression suite of already-fixed cases to catch whack-a-mole.
  2. Edit policy. Prefer general fixes to specific ones — this is the post's laundry-list-vs-canonical-examples point restated as bias-variance. A rule naming a specific input is high-variance; a rule naming a principle generalizes. If you cannot articulate the principle behind a patch, that is the signal you are fitting a sample.
  3. Budget. Cap prompt growth. Every N iterations, delete the oldest instructions and re-run evals — most will not be load-bearing. Prompts accrete monotonically unless you force a deletion step.
  4. Model. Re-test on a different or newer model. Instructions that help only one model are usually compensating for that model's quirk, not encoding your task.

“One of the most common failure modes we see is bloated tool sets that cover too much functionality or lead to ambiguous decision points about which tool to use.”

How can you detect and mitigate this issue?

Detection, MECE by signal type:

  1. Static — no traffic needed. Embed tool descriptions and flag high-similarity pairs. Run the post's heuristic as an actual exercise: two engineers independently label which tool each of ~50 realistic requests should use, then measure agreement.
  2. Behavioral — from traces. Per-tool call rate (never-called tools pay token rent every turn for nothing); tool error and retry rate; thrash patterns where tool A is followed by tool B on the same subgoal; definitions as a share of total context.
  3. Interventional. Ablate: remove a tool, re-run evals. The only signal that is causal.

Mitigation, MECE by action:

  1. Merge near-duplicates into one tool with an enum parameter.
  2. Delete anything below a usage floor.
  3. Tier — a small always-loaded core set, plus dynamic loading of the rest (tool search, or per-task MCP server loading).
  4. Disambiguate in the description — state explicitly when not to use it and which tool to use instead. Cheapest fix, usually the highest yield.

“However, teams will often stuff a laundry list of edge cases into a prompt in an attempt to articulate every possible rule the LLM should follow for a particular task. We do not recommend this. Instead, we recommend working to curate a set of diverse, canonical examples that effectively portray the expected behavior of the agent.”

One person’s laundry list is another person’s diverse set. How to distinguish?

[Post] No criterion offered. Genuine soft spot in the post.

[Inferred] The distinction is real and testable. Three tests, cheapest first:

  1. Generalization test. Does this example teach a rule applicable to inputs you haven't written down? Canonical examples cover a region of the behavior space; laundry-list items are point masses covering only themselves. If removing an item breaks only the exact case it names, it is laundry.
  2. Coverage-vs-count test. Map examples against the task's dimensions (input type × difficulty × output form). A diverse set is few examples spread wide. A laundry list is many examples clustered in one corner — usually whatever broke most recently.
  3. Ablation test. Remove each item, re-run evals. Decisive, and the only empirical one.

[Inferred] Two fast heuristics that usually settle it without running any of the above:

  • Provenance. Laundry lists grow append-only from incident response; canonical sets are curated top-down from a task taxonomy. If your examples section is a log of past bugs, it is a laundry list regardless of how varied it looks.
  • Form. Laundry-list items are exceptions ("if X, don't do Y"). Canonical examples are demonstrations (full input → output pairs).

“agents built with the “just in time” approach maintain lightweight identifiers (file paths, stored queries, web links, etc.) and use these references to dynamically load data into context at runtime using tools”

What does it mean to “maintain lightweight identifiers”? Where and how?

[Inferred] "Maintain" covers three different storage locations, MECE by where the identifier sits:

  1. In the context window, as plain text. The agent ran ls or grep -l and the paths are now in the message history. That is the maintenance — a path costs ~10 tokens, the file ~10,000. No machinery required. This is most of what Claude Code does.
  2. In the environment, re-derived on demand. The identifiers are not held at all; the filesystem is the index, and the agent regenerates paths with glob/grep whenever needed. Zero context cost, one round trip.
  3. In external persistent storage. A NOTES.md, a to-do list, or the memory-tool file system the post introduces later. The durable form — and the only one that survives compaction.

[Inferred] The honest answer: this is mostly not a data structure you build. It is a discipline enforced in your tool contracts — make search return paths + line numbers + a two-line snippet rather than file contents, and the identifiers maintain themselves. The engineering work is in tool return shapes, not in an identifier registry.

“primitives like glob and grep allow it to navigate its environment and retrieve files just-in-time”

These feel like very rudimentary ways to access relevant files? Is lexical similarity actually sufficient?

[Inferred] For code, the objection is weaker than it looks, for three reasons:

  1. Code isn't prose. Identifiers are exact, near-unique strings. Want UserAuthService? Grep it — every definition and call site, perfect precision. Embeddings are worse here: they blur exact tokens, which is the opposite of what you want.
  2. It's agentic, not single-shot. The comparison is not one grep vs. one embedding query. It is one embedding query vs. fifteen greps with a model reading results and reformulating between each. The loop supplies the semantics; grep supplies the lookup. Recall misses get recovered in later iterations.
  3. Freshness and zero infrastructure. No index to build, invalidate, or sync. On a repo the user is actively editing, a stale index is worse than none.

[Inferred] Where it genuinely fails: conceptual queries with no shared vocabulary — "where do we handle rate limiting?" when the code calls it ThrottleGuard. The model's world knowledge partly covers this (guess synonyms, grep again) but not reliably. That is the case for hybrid: semantic search over docstrings or file summaries to find the neighborhood, then grep for precision inside it.

“model to summarize and compress the most critical details”

How does the model know what to include in the summary and what the most important details are? Do you just ask it to figure that out itself or is it more sophisticated than that?

[Post] Answers more than you would expect. It names the categories — the model "preserves architectural decisions, unresolved bugs, and implementation details while discarding redundant tool outputs or messages" — and says engineers should be "carefully tuning your prompt on complex agent traces." So: it is a prompt, and the categories come from the engineer, not the model.

[Inferred] But it is more structured than "summarize this":

  1. Schema-driven, not free-form. Production compaction prompts specify an output skeleton — goal, decisions made, current state, open problems, files touched, next step. The model fills slots. Far more reliable than open-ended summarization.
  2. Deterministic pre-filtering runs first. Tool result clearing (the post's own example, now a Claude Developer Platform feature) involves no model judgment at all. Cheap deletions happen before expensive summarization.
  3. Recency passthrough. Recent turns are copied verbatim rather than summarized. The five-files detail is this same idea applied to files.

[Inferred] What is not happening: there is no training signal on "did this summary preserve what mattered." Importance is hand-specified and tuned offline against traces — which is exactly why the post frames it as a recall-then-precision prompt engineering exercise rather than a learned component.

“The agent can then continue with this compressed context plus the five most recently accessed files”

Why five? What happens if the files are very big?

[Inferred] On five — almost certainly an empirical constant, not a derived one:

  • It is a recency heuristic proxying for relevance. In coding, files you just touched are the ones you are about to touch. Cheap and usually right.
  • Five is roughly where marginal value crosses marginal token cost. At ~500–2,000 tokens per source file, five is ~2,500–10,000 tokens — low single-digit percent of a 200k window. Twenty files would be a meaningful fraction of a freshly cleared window.
  • It is the kind of number tuned once on internal traces and then frozen.

On big files — the more interesting question, and the post ducks it. Three possibilities:

  1. A token budget binds, not the count. Any sane implementation caps total tokens, so "five files" is really "up to five, subject to a budget." Otherwise one vendored 500k-line file blows the fresh window immediately.
  2. Truncation or re-read-on-demand. Either partial inclusion, or keep only the path and let the agent re-read the relevant span with offset/limit — which is precisely the just-in-time pattern the post argues for earlier.
  3. If neither: compaction becomes self-defeating — you summarize to free space, then immediately refill it.

[Inferred] Worth flagging the tension: re-injecting file contents is pre-loading, which contradicts the post's own just-in-time thesis. Re-injecting paths would be the consistent choice. The likely justification for pre-loading anyway: after a reset the agent has lost the reasoning that made those files relevant, so paying tokens to restore working state beats a slow rediscovery loop.

“Start by maximizing recall to ensure your compaction prompt captures every relevant piece of information from the trace, then iterate to improve precision by eliminating superfluous content.”

How would recall and precision actually be measured in this situation? Or is it more qualitative?

[Inferred] Both are measurable, but only once you fix a "unit of information" — which is the whole difficulty. Three levels, MECE by cost:

  1. Qualitative human review — what most teams actually do. An engineer reads a long trace, lists what matters, and checks the summary against it. Recall = fraction of that list present. Subjective and slow, but catches the failures that matter.
  2. Proposition-level with an LLM judge — the practical middle. Decompose the trace into atomic claims ("the auth bug is unfixed", "we chose Postgres over Mongo"). A judge checks each against the summary. Recall = claims preserved / claims extracted. Precision = summary claims supported by the trace / total summary claims — note this catches hallucination, not just verbosity.
  3. Downstream task accuracy — the only ground truth. Generate question-answer pairs from the full trace, then ask them of an agent given only the summary. Recall failure = questions it can no longer answer. This measures the thing you actually care about (can the agent continue) and is the standard formulation in long-context and memory benchmarks.

[Inferred] One precision note: the post's "precision" is not the classification-theoretic one. It means token efficiency — no superfluous content. So the operational metric is closer to compression ratio at fixed recall.

“the agent regularly writes notes persisted to memory outside of the context window”

  1. How does the model know what to include in the notes? Is it application-dependent?
  2. Why written to memory outside of context window? How does it know when to retrieve?

(1) What to include

[Post] Says both things, in tension. Engineered on one hand (Claude Code's to-do list, a NOTES.md convention); emergent on the other — the Pokémon agent, "without any prompting about memory structure," develops maps of explored regions, tracks achievements, and maintains combat strategy notes.

[Inferred] Reconciliation: the trigger and affordance are engineered (a tool exists, the prompt says to use it); the schema is often emergent. And yes, strongly application-dependent — what to persist is whatever must survive a reset in that domain: architectural decisions for coding, entity relationships for research, map/inventory/strategy state for a game.

(2) Why outside the window, and when to retrieve

Why outside: [Post] "provides persistent memory with minimal overhead" — notes survive compaction and resets; in-context content by definition does not.

[Inferred] A second reason the post does not give: notes are rewritable. Context is append-only — you cannot revise an earlier message, only append a correction, and both then sit in the window. A file can be edited in place, so stale state gets replaced instead of accumulating alongside its correction.

When to retrieve: [Post] Doesn't say, beyond "these notes get pulled back into the context window at later times" and "after context resets, the agent reads its own notes."

[Inferred] Three mechanisms in practice, MECE by who decides:

  1. Unconditional on reset — the post-compaction prompt says "read your notes file."
  2. Model-decided — the memory tool is a filesystem; the agent reads when it judges it needs to. Unreliable without prompting.
  3. Always-on injection — small files (a to-do list) get re-rendered into context every turn.

[Inferred] The real gap: retrieval timing is the weakest link in the whole scheme. An agent that does not know it forgot something has no trigger to go look.

“After context resets” What does reset mean here?

[Post] Undefined. Used once, in the Pokémon section, next to "coherence across summarization steps" — implying reset ≈ compaction event.

[Inferred] It conflates three things worth keeping separate, MECE by what survives in-band:

  1. Compaction. The window is replaced by a summary of itself. Some information survives in-band. This is what the post means in that sentence.
  2. Hard clear. The window is wiped to system prompt plus a fresh start; nothing survives in-band, only external files carry over. This is what "reset" most literally means, and what a game harness does when it hits the limit.
  3. New session. A separate process started later — the same information situation as (2) but across time. This is what the memory tool targets: "maintain project state across sessions."

[Inferred] Why the distinction matters: the note-taking argument is only load-bearing for (2) and (3). For (1), a well-tuned compaction prompt can carry state in-band. Notes still win because they are lossless for what they cover and do not degrade across repeated resets — a summary of a summary of a summary drifts; a file does not.

How would you improve this system if you were building it today?

Ordered by expected value:

  1. Make context management measurable, not advisory. The post's entire toolkit is heuristic with no evaluation story. Build a continuation benchmark: sample real traces, compact them, then measure whether the agent can answer trace-derived questions and complete the task. That converts "tune your compaction prompt" into a regression test.
  2. Tiered editable state instead of one summarization pass. Compaction today is monolithic and lossy in a single shot, and re-summarizing a summary compounds drift. Better: verbatim recent turns, plus a structured state document (decisions, open problems) that is edited rather than re-summarized, plus cold storage that is searchable but unloaded.
  3. Treat tool definitions as retrievable context. MCP made bloated tool sets the default. Load a small core set and retrieve the rest on demand via tool search. This directly attacks the failure mode the post names but does not solve.
  4. Give note-taking an index. Per Q16, retrieval timing is the weak link. Keep a cheap always-in-context table of contents of what is in memory, so the agent knows that something exists before deciding whether to read it. Knowing you forgot is the precondition for looking.
  5. Reconcile just-in-time with the five-files behavior. Post-compaction, re-inject file paths plus one-line summaries, not contents. Cheaper and consistent with the rest of the argument.
  6. Measure the write side, not just the read side. Everything in the post is about what enters context. Nobody measures what the agent writes — note quality, note staleness, contradictory notes accumulating. That is where long-running agents actually fail.
2026 Blogpost original

How Abnormal Taught AI Agents To Write Detectors

My notes

What I didn’t understand or am unsure about

Question

Answer

“customer-reported misses in Detection 360”

  1. What is Detection 360?
  2. How do customers realize misses happened?

(1)

An in-product submission channel: "Submit a missed attack or a false positive incident from within the product's Detection 360 tab" (navigation: Investigate → Detection 360). The case object carries sender, recipient, subject, submitted-by, case number, status (Submitted/Resolved), and a VIP flag. The platform page summarizes the loop: "Detection 360 investigates, explains the outcome, updates the models." Case lifecycle runs "from initial classification through investigation, remediation, and deployed detections." Turnaround is marketed as "hours, not days," and the loop reportedly got "94% faster."

(2)

Abnormal only ever uses analyst-initiated framing ("An analyst reports a vendor impersonation attack"). Their AI Security Mailbox "connects to existing phishing-report buttons, abuse mailboxes, G Suite alerts, and third-party tools" — but no page states that end-user reports automatically create Detection 360 submissions. They are documented as separate paths.

[Inferred] The realistic channels, MECE by who notices:

  1. End user reports it — phishing button or abuse mailbox, then a SOC analyst escalates
  2. SOC proactive hunting — searching mail for indicators pulled from threat intelligence
  3. A different tool catches it downstream — EDR fires on the payload, web proxy blocks the link, a SEG flags a sibling message
  4. Post-incident, working backwards — someone clicked, credentials were phished, money moved
  5. External notification — the impersonated vendor, a partner, or law enforcement
  6. Purple team or phishing simulation — a deliberate test that got through

Where does this agentic pipeline fit into the broader system? Given latency constraints

It is clearly offline and asynchronous. The post gives four tells:

  • It is triggered by a submission, not by an arriving message
  • It evaluates "across a larger slice of the customer's real traffic" — historical replay
  • It loops "multiple times"
  • End-to-end is "hours"

So the pipeline sits beside the detection engine, not inside it. Same architecture as the 2024 rules post: expensive LLM work happens once per attack pattern offline; a cheap deterministic artifact runs online.

The three-tier picture:

Tier

Cadence

What runs

Budget

Online scoring

every message

ensemble plus rules/detectors

sub-second [Search]

Agentic detector generation

per Detection 360 submission

Deep Research plus Adaptive Detection agents, LLM calls, replay eval

hours [Post]

Core model retrain

weekly

full retrain

days to weeks [Search]

[Inferred] This places the agentic loop in the middle speed tier — faster than retraining, far slower than a scoring decision. Which is exactly the gap the 2024 post identified: "adding new [features] involves retraining the model, which takes weeks." The agents fill that hole. That is the one-sentence answer to how it fits.

“how they normally communicate”

In what form?

Sender-history profile, computed per sender-recipient pair and per sender-tenant:

  • Has this sender written to this recipient before, and how often
  • Typical send times and cadence
  • Typical infrastructure — sending IPs, domains, authentication posture
  • Typical recipients within the organization
  • Typical subject/topic distribution and tone
  • Signature blocks and formatting conventions

Form: almost certainly a structured summary — counts, rates, first-seen dates — plus probably a handful of retrieved prior messages for the agent to read.

“how the message compares to typical organizational patterns”

In what form?

  • Does this organization receive Teams event invites at all? From which platforms?
  • Which vendors are established relationships versus first contact?
  • What is the normal distribution of external senders, of financial-request language, of invite mail?

“what signals point toward intent or impersonation”

In what form?

[Post] Not specified, but the example output enumerates a real subset: sender domain and its subdomain lineage, authentication results (SPF/DKIM), body link destinations, subject content, delivery mechanism. [Post] elsewhere: "thousands of rich signals."

[Inferred] The likely full set adds display-name-vs-address mismatch, first-time-sender flags, phone numbers extracted from the body, brand mentions inconsistent with sender identity, and the text-model scores from the NLP post (urgency, tone).

[Inferred] Form: most likely a text serialization of the feature record — name/value pairs with short natural-language annotations — because the agent's output reasons over them in prose. Standard pattern: render structured features into a prompt, let the model reason, give it tools to query more.

[Inferred] The design question worth raising: with "thousands" of signals you cannot render them all into a prompt. So something selects. Either the agent has tools to query attributes on demand — which the post hints at, "its tools let it verify those intuitions against real data" — or a retrieval step pre-selects. On-demand querying is the better design and is what the post implies. Same feature-selection-as-retrieval problem flagged in Q4 on the 2024 post; it recurs here.

“This is a **trusted platform abuse** attack - using Microsoft Teams' event registration system to deliver phishing content.“

If the Deep Research Agent already determined it’s an attack, what’s the Adaptive Detection Agent doing?

[Post] The Deep Research Agent's output is "a detailed behavioral profile of the attack" — prose, explaining why this one message is malicious.

[Post] The Adaptive Detection Agent's job is to "find the attributes that represent the underlying behavior behind the attack," check statistical and semantic relevance, apply second-order abstraction, test against traffic, and iterate.

Three ways to state the difference:

  1. Explanation vs. classifier. The DRA answers "why is this message malicious?" The ADA answers "what decision rule separates this class of message from all normal traffic?" A perfect explanation may cite features that do not generalize, or that are not even available as production attributes.
  2. Prose vs. an executable artifact grounded in the production feature space. The DRA can say "uses legitimate Microsoft Teams infrastructure." The ADA must translate that into attributes the engine actually computes at scale, in sub-second, and verify each one separates on real traffic. The DRA has no obligation to check whether "uses Teams infrastructure" also fires on ten thousand legitimate meeting invites a day. The ADA does.
  3. One message vs. a population. The DRA reasons about a single known-bad message with zero false-positive cost. The ADA operates against a one-in-a-million FP budget over a traffic distribution.

[Inferred] The analogy: the Deep Research Agent writes the incident post-mortem; the Adaptive Detection Agent writes the regression test. Knowing exactly why the bug happened does not hand you the test that catches the next one.

[Inferred] One more asymmetry: the DRA's determination is retrospective, on a message already known to be a miss. The ADA must produce something that works prospectively, on messages nobody has labeled.

“The problem is that without any deeper understanding, teams end up with rules that are technically correct but practically useless. A system might flag that 100% of attacks in a cluster lack an attachment, which is true but meaningless since most legitimate emails don't have attachments either. Statistical relevance alone produces a lot of noise without the next critical piece”

This seems like it’s only true for the most rudimentary definition of statistical relevance?

Yes. The critique is correct, and this is the strongest pushback available on this post.

Their example is a statement about P(no attachment | attack) — recall on the positive class alone, with no reference to the negative class. Every standard criterion rejects it immediately:

Criterion

Value for their example

Verdict

Precision, P(attack given no attachment)

approximately the base rate, so near 0

rejects

Lift, P(no attachment given attack) / P(no attachment)

approximately 1

rejects

Mutual information / information gain

near 0 (near-independent)

rejects

KL divergence between class-conditional distributions

near 0

rejects

So what they describe is not "statistical relevance" — it is support on the positive class. A first-year decision-tree splitting criterion catches this. They picked a weak example and undersold their own argument.

[Inferred] The real case for semantics is different, and stronger. Three genuine reasons statistics alone is insufficient here:

  1. Multiple comparisons at tiny n. A submission yields maybe a handful of attack messages against thousands of candidate attributes. At small n and high dimensionality, many attributes separate perfectly by chance. Statistics cannot distinguish causal from coincidental at that sample size. Semantics functions as a prior over hypotheses that prunes the search space before testing — exactly what you need when the multiple-testing burden is otherwise fatal. Note this is a well-posed statistical argument, not an anti-statistical one.
  2. Durability under adversarial shift. Statistical relevance is measured on the past and says nothing about which correlations survive the attacker's next move. "Sender domain equals X" is maximally statistically relevant and minimally durable. Semantics is how you predict stability — their own second-order-thinking point, and the argument they should have led with.
  3. Cost asymmetry is not in the statistic. At a base rate under 1 in 10 million and an FP target of 1 in a million, an attribute with excellent in-sample lift can still be catastrophic at production volume. Understanding what an attribute means is how you anticipate where it fires outside your sample.

[Inferred] Accurate framing: semantics is a prior and a durability filter, not a substitute for statistics.

[Inferred] The charitable reading, in their defense: at n of about 5 attack messages, most rigorous statistical criteria are genuinely underpowered — you cannot estimate lift reliably from five examples. So their practical situation may justify leaning hard on priors. That is probably what they meant; the attachment example just does not communicate it.

“It has enough statistical sense”

How? Reasoning, tool-use or something else?

[Inferred] Tool use, not reasoning. The LLM does not compute statistics in-context. It calls something that queries traffic and returns counts and rates, then reasons over the returned numbers. That is the only design that works, for three reasons:

  1. LLMs are unreliable at arithmetic over large numbers, and these are counts over millions of messages
  2. The data is not in the context window — it is in their feature store and data lake
  3. The post explicitly describes an iterate-and-verify loop, which requires an external oracle

[Inferred] The likely tool surface: "for attribute A and recipient R, what fraction of the attack messages have value v, and what fraction of normal traffic does?" plus "run this candidate detector against this traffic slice and return the hits."

The "statistical sense" that actually lives in the model is choosing which query to run and interpreting the result — not computing it.

[Search] Consistent with their broader practice: their backtesting infrastructure already does exactly this kind of replay, "re-hydrated exactly as they would have been during online scoring."

[Inferred] Worth naming the pattern: LLM-as-orchestrator over a statistical oracle. Same division of labor as the 2024 post — LLM does translation and hypothesis generation, deterministic software does evaluation. The same team making the same architectural choice twice, two years apart, is a real signal about their design philosophy.

“The second-order approach is to ask: what's true about the domain that makes it suspicious? Is it a domain the organization has never seen before? Is it newly registered? Those characteristics will still hold when the attacker switches to a different domain next time.”

Is the Adaptive agent simply trying to answer yes or no to these questions, or use them in its downstream determination in the same iteration as well?

[Post] answers this indirectly: "its tools let it verify those intuitions against real data before committing to a rule." So the sequence is hypothesize, verify, commit — and the verified abstraction becomes a predicate in the rule, not a fact that is consulted and discarded.

[Post] The end state confirms it: "The result is detectors that generalize well beyond the specific messages they were built from." That only holds if the abstraction (domain novelty, domain age) is what the shipped detector tests, not the raw domain.

[Inferred] The flow for the PayPal example:

  1. Observe — all the attacks share sender domain *.teams-events.com
  2. Hypothesize the second-order property — "a domain this organization has never seen before," and/or "recently registered"
  3. Verify against traffic — does never_seen_domain actually separate these attacks from this recipient's normal mail?
  4. Commit — if yes, the abstraction goes into the detector; the raw domain does not

So: both. It is a hypothesis that gets answered yes/no and, if yes, becomes part of the rule in that same iteration. The yes/no is the verification gate; the attribute is the payload.

[Inferred] Worth naming: this is generate-and-test over a hypothesis space — classic program synthesis with an execution oracle. The LLM proposes candidate abstractions; traffic data accepts or rejects them.

“they latch onto the most distinctive surface features of the example they were shown”

What does distinctive mean here?

[Inferred] Precisely, it means maximally identifying of the specific example, relative to the corpus prior — the lowest-probability, highest-information strings in the input. For the PayPal message that is the exact domain jorjarounsevell.s08.usa1.teams-events.com, the exact phone number, the exact subject line.

Three framings of why models gravitate there — same phenomenon, different vocabulary:

  1. Information-theoretic. A rare string carries high self-information. Asked "what characterizes this message," the rare strings are the highest-signal answer to that question as literally posed.
  2. Learning-theoretic. From n of 1, the minimum-description-length hypothesis that perfectly separates the example is the example itself. Memorization is the optimal fit. Generalization requires a prior preferring abstract hypotheses, and nothing supplies that prior unless you put it there deliberately.
  3. Shortcut learning. The model takes the feature with the highest in-sample discriminative power at the lowest reasoning cost. Exact string matching is both.

[Inferred] The essential point: distinctive is measured on the example; predictive must be measured on the class. Those coincide only if the example is representative in the right way — and for an adversary-generated example they systematically do not, because the attacker controls the distinctive features precisely so they can rotate them.

So in this domain, "most distinctive" and "most durable" are close to anti-correlated. That inversion is the entire reason second-order thinking has to be engineered in rather than emerging naturally.

“We worked to guide the agent toward behavioral abstractions over raw values”

By prompting, fine-tuning or what?

[Inferred] The wording points away from fine-tuning and toward three softer levers, in order of likely impact:

1. Attribute vocabulary design — probably the biggest lever, and structural rather than persuasive. If the attributes exposed to the agent are already behavioral — never_seen_sender, from_fqdn_age_in_days, both real names from their 2024 DSL — then a first-order rule becomes hard to express. You do not have to talk a model out of matching a raw domain if no attribute matches raw domains. This is guidance by construction, and it is far more reliable than instruction.

2. Prompting. Explicit instruction plus few-shot examples contrasting first-order and second-order rules for the same attack. Cheap — and "we worked to guide" / "something we spent a lot of time tweaking" is precisely what iterating on a prompt sounds like.

3. The verification loop as an implicit filter. A first-order rule passes local eval but generalizes badly in the broad eval — it will not catch other instances of the attack type. [Post] treats that as an explicit signal: "Did it find other instances of this attack type that were previously missed? If so, that's a strong signal." So the evaluation selects for generalization even when the model's first instinct is first-order. Guidance via reward rather than instruction.

[Inferred] Why probably not fine-tuning: you would need a labeled corpus of good-vs-bad detectors, which is exactly the scarce thing this system exists to produce; "guide" and "tweaking" are not fine-tuning vocabulary; and [Search] their documented fine-tuning work (7B Mistral/Llama-2 via LoRA) is described for detection and classification, not agent behavior.

“First, the agent runs a fast local evaluation. It takes the attack messages it already identified along with a random sample of normal traffic for the recipient and tests the detector against them.”

The random sample would need to be huge for this to work, right?

Yes — and the post effectively concedes it, twice.

[Post] The local stage is explicitly not a precision estimate: "This is intentionally lightweight. The goal is to iterate quickly." And the broad stage exists because the sample is insufficient: "it lets us catch any false positives hiding in the broader traffic that didn't show up in the smaller sample."

So the local sample's job is screening, not estimation — kill obviously broken candidates cheaply before paying for a full replay.

[Inferred] The arithmetic, worth having ready. Observing zero false positives on a sample of size n bounds the true rate only weakly. By the rule of three, with zero events in n trials the upper bound of the 95% confidence interval is about 3/n:

Clean sample size

95% upper bound on FP rate

vs. the 1-in-a-million target

1,000

about 3 in 1,000

3,000 times worse

100,000

about 3 in 100,000

30 times worse

3,000,000

about 1 in 1,000,000

at target

A "random sample of normal traffic for the recipient" — one mailbox, days or weeks of mail, so hundreds to low thousands of messages — is three to four orders of magnitude short of supporting a precision claim. It is entirely adequate for what it actually does.

[Inferred] The second, subtler problem: the sample is drawn from the wrong distribution. Randomly sampling one recipient's normal mail gives mostly easy negatives — internal mail, known vendors, newsletters. The false positives that matter are near misses: legitimate Teams invites, real PayPal notices, genuine first-time vendor contact. Random sampling surfaces those only in proportion to their frequency, which is low. So you systematically fail to detect the exact failure mode the detector is most likely to have.

[Inferred] The fix: stratified hard-negative sampling. The Deep Research Agent has already characterized the attack as "Teams event registration abuse." Use that to sample negatives conditioned on the characterization — real Teams event invites, real event-platform mail, real PayPal correspondence. Small, targeted, far more informative per message than a random draw. Standard hard-negative mining, and it fits their architecture cleanly since the characterization is already in hand.

“Every message the detector would have flagged gets reviewed, and we LLM-label each one to determine what it actually is.”

If we’re using an LLM to label it again, wouldn’t that just lead to artificially inflated agreement?

Good instinct, and partly right — but the failure mode is narrower than self-grading.

First, why it is not circular in the obvious way: [Post] the detector is not an LLM. It is a rule over attributes; the LLM wrote it offline. So at evaluation time you have a symbolic classifier being graded by a language model — different mechanisms, not a model scoring its own forward pass. That defuses the strongest version of the objection.

[Inferred] But three genuine correlation risks remain, MECE by where the leak enters:

  1. Shared priors between writer and grader. If the same model family that reasoned "Teams invite plus urgency equals attack" also labels flagged messages, it will tend to agree — not because the detector is right, but because both inherit the same notion of what phishing looks like. Legitimate mail that merely looks phishy to that model gets labeled malicious, and a real false positive is scored as a true positive.
  2. Context contamination. If the labeler sees the detector's rationale or the Deep Research Agent's attack profile, it is being primed with the hypothesis it is supposed to test. Strongest version of the problem, and the most easily avoided.
  3. No ground-truth anchor. [Post] gives no indication that LLM labels are ever calibrated against human labels. Without that you cannot distinguish "the detector is precise" from "the labeler agrees with the detector." The number is uninterpretable in isolation.

[Inferred] What to do, in order of return:

  • Blind the labeler — it sees the message and nothing about the detector or the attack profile. Cheap, eliminates risk 2 outright.
  • Use a different model family for labeling than for generation — decorrelates errors. Same logic as their multi-tokenizer ensemble in the NLP post, applied to judges.
  • Anchor on a human-labeled calibration set — periodically measure the labeler's own precision and recall against expert labels, then report detector precision corrected for labeler error.
  • Human review on the tail — auto-accept where the labeler is confident and volume impact is small; escalate high-volume or VIP-recipient detectors.

[Search] They already know how to do this. Their 2024 post Testing GenAI Products describes using "a separate judge LLM, with access to our accuracy rubric" and running "separate 'unrestricted' red team LLMs." The judge-LLM discipline is established internally — this 2026 post just does not say whether those safeguards apply here.

“The agent loops through this cycle multiple times, refining the detector with each pass until it can reliably catch the attack without compromising on safe emails.”

Stopping criteria?

[Inferred] What a real implementation needs, MECE by termination reason:

  1. Success — precision on the broad eval above a threshold, and recall on the submitted attacks at or near 100%, and firing volume within budget. All three, not just the first.
  2. Budget exhaustion — max iterations, or max LLM/eval spend. Necessary because the loop is not guaranteed to converge.
  3. Plateau — no improvement over k consecutive iterations. Cheaper than burning the full budget.
  4. Abandonment — declare the attack not separable with the available attributes and escalate to a human. [Post] never mentions this possibility, and it must exist. Some attacks genuinely cannot be caught by a rule over current attributes; the honest outcome is "this needs a new feature" or "this needs the model," not a marginal rule. A loop with no failure exit will eventually ship something that satisfies the letter of the criteria and not the spirit.

[Inferred] The real risk the post never names: eval-set overfitting through iteration.

Every loop uses evaluation results to modify the detector. That is adaptive data analysis — after k rounds the measured precision on that traffic slice is optimistically biased, and the bias grows with k. You are effectively doing gradient descent on your test set.

Mitigations, strongest first:

  • A final held-out gate — a fresh traffic slice never seen during iteration, evaluated exactly once. Fail there and the detector does not ship, regardless of how good the iteration numbers looked.
  • Fresh slice per iteration, so no single slice gets fit repeatedly.
  • Penalize iteration count — a detector that needed 12 rounds is more suspect than one that passed in 2, because it has absorbed more evaluation information.

This connects to Q32: the two-stage eval already has the right shape; it needs a third, sacrosanct stage that iteration never touches.

“Every submission kicks off the full loop”

What is a submission? Does the loop get kicked off by a single false negative?

[Search] A submission is a Detection 360 case, filed in-product: "Submit a missed attack or a false positive incident from within the product's Detection 360 tab." The case object carries sender, recipient, subject, submitted-by, case number, status, and a VIP flag.

So yes — a single customer-reported message can start the loop. Note it can be a false positive as well as a false negative, though this post only describes the missed-attack path.

[Inferred] But the loop does not stay at n of 1, and the post quietly tells you so. It says the agent tests against "the attack messages it already identified" — plural. That means between submission and generation, one reported message gets expanded into a cluster: the Deep Research Agent (or something upstream) finds related messages in the customer's traffic. This has to happen, because you cannot do statistical attribute selection on a single example.

[Inferred] Two consequences:

  1. Cluster quality gates everything downstream. Expand too narrowly and you are back to n of 1, with no evidence to verify the second-order abstraction against. Expand too broadly and you are characterizing a mixture, so the detector ends up loose or incoherent. The post says nothing about how this expansion works, and it is arguably the most load-bearing undocumented step in the pipeline.
  2. A single submission is an unvalidated trigger. [Search] No public Abnormal source addresses adversarial or poisoned submissions. An attacker with access to a tenant's Detection 360 — or who can get a legitimate message reported as an attack — could induce the system to auto-generate and deploy a detector that fires on legitimate traffic. [Post] There is explicitly no human in the loop: "No analyst queue, no manual review, no waiting." The evaluation stages are the only defense, and they are graded by an LLM (see Q33).

How would you improve this system if you were building it today?

Stage 1 — Understand the attack

  • Fix the input bias, or at least measure it. Detection 360 misses are the misses somebody noticed, not a random sample (Q21). Instrument it: periodically sample random traffic for expert labeling to estimate the unreported miss rate, so you know how much of the problem the loop can even see. Without this you cannot tell whether "hours to protection" covers 10% or 90% of your misses.
  • Document and harden cluster expansion (Q36). Report cluster size and purity alongside every generated detector, at minimum.

Stage 2 — Generate the detector

  • Enforce second-order by construction, not by prompting (Q31). Expose only behavioral attributes, or type-check candidate rules and reject raw-literal predicates. Guidance that lives in the schema cannot be talked out of; guidance that lives in a prompt can.
  • Generate a population, not a single detector. Sample K candidates, evaluate all, keep the precision/recall Pareto front. More LLM calls, but calls are cheap relative to a bad detector.

Stage 3 — Evaluation

  • Hard-negative sampling in the local stage (Q32), conditioned on the attack characterization the Deep Research Agent already produced.
  • Blind the LLM labeler, use a different model family, add a human-labeled calibration set (Q33) so labeler error can be corrected for rather than ignored.
  • A held-out final gate the iteration loop never touches (Q35). The change to argue hardest for — without it, the reported precision is optimistically biased by construction.

Stage 4 — Precision optimization

  • Explicit stopping and abandonment criteria (Q35), including the option to fail and escalate rather than ship a marginal detector.
  • Penalize iteration count as a proxy for overfitting risk.

Stage 5 — Deployment

  • Blast-radius controls. Auto-deploy in score-contributing mode first; promote to verdict-forcing after a live observation window. Cap how much a newly generated detector can quarantine per hour. Auto-rollback on FP-rate breach.
  • Tiered human review. Fully automatic for low-volume, low-impact detectors; mandatory review above a volume or VIP-recipient threshold. Preserves the "hours, not days" claim for the common case while capping tail risk — and that framing is how to sell it internally without gutting the product story.

Cross-cutting

  • Lifecycle and decay. Same gap as the 2024 post, but worse here, because generation is now fast enough to accumulate rules quickly. Per-detector monitoring (fire rate, precision, and marginal catch — messages it catches that nothing else does), automatic demotion when a detector stops catching, expiry by default.
  • Reconcile per-tenant vs. global (Q21). There should be an explicit promotion path: a tenant-local detector that proves out across multiple tenants gets promoted to global, at a higher evidence bar.
  • Submission trust model (Q36). Provenance checks, rate limits, and a poisoning red-team exercise. [Search] They already run red teams for prompt injection on their GenAI products; this loop deserves the same and there is no public evidence it gets it.
  • Feed detectors back into the model. Detectors are a fast patch for a slow retrain. Mine the accumulated detector population for recurring attribute combinations — those are missing features for the core model. Over time the detector layer should shrink as the model absorbs what it keeps rediscovering. Without that, you are building an ever-growing rule layer and calling it progress.
2025 Blogpost original

How Ramp Fixes Merchant Matches with AI

The existing transaction-merchant mapping review system was slow and created additional workload for customers and Ramp support. This is a problem now and is likely to become a much bigger problem as Ramp grows.

My notes

Summary

Context

Ramp customers want to be able to match transactions to merchants because it enables them to analyze spending by merchants, set up merchant restricted funds etc. Ramp allows users to submit requests to fix merchant classifications by providing a new merchant name, website, and category. These requests used to be reviewed by a combination of customer support, finance, and engineering teams.

Problem

The existing transaction-merchant mapping review system was slow and created additional workload for customers and Ramp support. This is a problem now and is likely to become a much bigger problem as Ramp grows.

Solution

Agentic approach to transaction-merchant mapping.

Some key insights/new things I learned

  1. For each transaction, a payment processor like Stripe sends transaction metadata in the card acceptor. This metadata includes the name of the business the user transacted with (although the format for this can be unusual) and a Merchant Category Code (MCC) which is supposed to describe the services the business does (but can be misleading for businesses that offer multiple services).
  2. A single merchant usually has multiple card acceptors.
  3. Merchant mapping can be difficult because the card acceptor names can be vague or outdated.
  4. Pipeline
    1. LLM reviews the user request and takes actions from a pre-specified list.
    2. Input/Features
      1. Transaction card acceptor name and MCC.
      2. Extracted merchant names, addresses, and line items from related receipt images.
      3. User-provided memos for related transactions.
    3. Retrieve related merchants to identify correct action
      1. Hybrid search - semantic similarity search using transaction embeddings and lexical similarity the “requested name” field provided by users.
    4. Actions
      1. Create a new Ramp merchant.
      2. Update an existing merchant.
      3. Reassign the transaction to a more fitting merchant.
    5. Guardrails
      1. Require that the LLM always chooses one of the provided actions
      2. If the LLM chooses to reassign a transaction to another merchant, the target must be in the supplied list
      3. If the LLM hallucinates, we inform it of its mistake and have it retry until we get a valid response
  5. Evaluation
    1. First, they did a manual review of LLM’s responses on a few select users and transactions.
    2. As it was rolled out to more users, track % of user quests that lacked followup requests (they assume this means the LLM acted properly on the initial request).
    3. Request rejection rate
    4. LLM as a judge
      1. In the event of a change, the LLM judge determines whether our agent's LLM action resulted in an improvement according to the user's request.
      2. In the event of a rejection, the LLM judge determines whether the rejection was reasonable.
    5. They also ran a shadow mode evaluation to see how the agent would behave on customers' transactions before actually rolling it out to those customers.

What I didn’t understand or am unsure about

Question

Answer

Do different payment processors send different kinds of data? If yes, would this affect the approach you use?

Yes. Ramp likely has a "normalization" or "ETL" (Extract, Transform, Load) layer that maps various processor formats into a standardized internal schema before the AI agent ever sees it.

What format do they use for the input to the LLM?

Not stated in the blogpost but probably something along the following lines:

  1. System Prompt with Guardrails
  2. Transaction Metadata, User Request Data, Retrieved merchants, specified actions provided as a structured list like JSON or Markdown

Why use an LLM instead of a simpler model?

Simpler models

  1. can’t provide explanations as easily
  2. are harder to use for multimodal and unstructured data
  3. would struggle with zero-shot because they don’t have world knowledge from pre-training

“By bringing in merchants using these two strategies, the LLM can inspect the names, websites, and categories of a good selection of merchants. It can then decide whether to create a new merchant or modify existing ones.”

It doesn’t look like they’ve mentioned how to produce a final list for the LLM. What’s our best guess for how they did the final ranking from the two lists?

Simple approach

  1. Take k/2 from each list to get k total results

Cross-Encoder Re-ranker

  1. Pull n >> k from each list to get 2n total results
  2. Use a cross-encoder (using all the features you have for both the transaction and the merchant) to rerank and filter for the final k
  3. Have LLM pick from this final list

Did they use any kind of few-shot prompting? If not, is that likely to have been helpful?

Unclear from the blogpost whether they used it.

It could be helpful, especially for edge cases like the government/Service Fee stuff and making it clear to the LLM what kind of explanation you want to provide users. Examples of both correct and incorrect responses could be helpful.

“We require that the LLM always chooses one of the provided actions”

It doesn’t look like they’ve mentioned how they actually enforce this restriction. What’s our best guess for how they did this?

They say “If the LLM hallucinates, we inform it of its mistake and have it retry until we get a valid response.” so it looks like they have retry logic to help with this.

They could also use Structured Output (potentially through Constrained Decoding).

“If the LLM hallucinates, we inform it of its mistake and have it retry until we get a valid response.”

What guardrails could they have put in place to avoid getting stuck in infinite loops?

  1. Limit number of retries
  2. Dynamic temperature adjustment
  3. Meta-Prompting for Errors (provide explanation of failure to decrease likelihood of future failures in retries).

Once you hit the maximum number of retries without a valid response, you can escalate to humans.

“We take our primary LLM's rejection reasoning and then use a second LLM to rewrite it in plain language.”

Is the second LLM there because the primary LLMs rejection reasoning is unlikely to be very accessible to users? If yes, how could they have set up the second LLM to address that gap?

Yes, it’s better to exclude the human-optimized explanations from the primary LLM to avoid overloading it and keep it focused on its main job.

The second LLM can be prompted with something like

“You are a helpful Ramp Customer Support agent. Your goal is to explain a technical rejection to a user in a friendly, clear, and non-technical way. Do not use terms like 'MCC' or 'LLM' or 'Card Acceptor.' Use the provided evidence to explain the 'Why' simply.”

“It is also important to recognize LLMs' knowledge that has been distilled as part of their training.”

If this is the case, would it be helpful to replace the LLM every so often to incorporate this new knowledge? Or would a better strategy be to just store all instances of name changes and use a mapping between old names and new names?

Leveraging the User-Provided Data is probably a better way to update the LLM’s knowledge. You only need to do it once for each merchant whose name changed.

Storing mappings is complicated. Multiple companies can have the same name, only distinguishable using metadata like location, industry etc. Probably not worth it.

They mention “multimodal retrieval augmented generation (RAG)”.

Where is the multimodal part used?

Related receipt images; could use a Vision-Language Model (VLM) or a specialized OCR-to-text pipeline

How is the LLM as a judge supposed to know if the “agent's LLM action resulted in an improvement according to the user's request”?

The Judge is likely using a Rubric-Based Evaluation. The prompt for the Judge LLM for changes could look like this:

"Compare the Original Merchant with the New Merchant created by the Agent.

  1. Does the New Merchant better match the Receipt Evidence (Line items: Fuel)?
  2. Does the New Merchant fulfill the User's Intent (Requested: Gas Station)?
  3. Is the New Merchant a valid entity (not a hallucination)?

Decision: If 'Yes' to all, mark as IMPROVEMENT. If the agent rejected the request, check if the User's input was 'Placeholder' or 'Malicious'—if so, mark it as REASONABLE REJECTION."

They said using LLM as a judge “Requires Cycles”; what does this mean?

Not sure what “Cycles” refers to here but I guess the general idea is that it requires a lot more work to set up. Maybe some degree of iteration on the prompt, input data etc to get it working correctly.

What is the “Mean Daily Improvement Rate on Classification Correction Requests”?

The 98.8% figure means that on a typical day, the Judge finds that 98.8% of the agent's changes resulted in a more accurate merchant record than what existed before. It is essentially an "Accuracy Score" for the actions the AI chooses to execute.

The reasonable rejection rate seems really low (one-third of legitimate user feedback is being ignored)? Should they be concerned about this?

Probably a worthwhile risk given that accepting a bad user request will mess up the database.

How would you improve this system if you were building it today?

  1. Agentic Search for relevant merchants
    1. An agent could look at the merchant’s website, social media or Google Maps page to understand it better. Addresses from receipts can also be reverse-engineered to find Google Maps pages which can then provide a lot of additional useful data.
  2. Speculative Decoding or Router Architecture (easy requests handled by fast model, harder requests handled by better model) to speed up responses
  3. Consider adding some LLM as a judge negatives to the few-shot examples in the prompt
2026 Blogpost original

Rebuilding the Review Algorithm to Increase Accuracy and Speed

There are three problems that they're trying to solve in this blog post: The previous algorithm produced two fields per result: summary and additional context. This was confusing to users. Citations were not granular enough. Model…

My notes

Blogpost link: Rebuilding the Review Algorithm to Increase Accuracy and Speed

Important background: Speed Up Your Analyses with Review Tables in Harvey

Summary

Context

Harvey has a feature called review tables where each row represents a single document and each column represents an extracted feature for that document. This allows legal teams to extract information from large sets of documents at once

Problem

There are three problems that they're trying to solve in this blog post:

  1. The previous algorithm produced two fields per result: summary and additional context. This was confusing to users.
  2. Citations were not granular enough.
  3. Model reasoning was not available to users,

Solution

  1. Combined fields
  2. Built sentence-level citations
  3. Shared reasoning with users.

What I didn’t understand or am unsure about

Question

Answer

“citations were attached to the cell as a whole”

Does “attached to the cell” here mean that the citation mapping was (full generated text) <> (source document)? Or was it (sentence in generated text) <> (source document)?

I think they mean that the original was (full generated text) ↔ (source document) and they moved to (sentence in generated text) ↔ (source document).

“Previously, review tables generated sources through an algorithm that combined model-generated text and fuzzy matching.”

What would this methodology might have looked like concretely?

This could potentially look something like the following:

  1. Generation step: The model produces a free-form answer for the cell (e.g., a summary of a contract clause) without any grounding mechanism built into generation itself.
  2. Post-hoc matching step: Separately, the system takes that generated text (or key phrases from it) and runs a fuzzy string-matching algorithm (e.g., something like Levenshtein distance, n-gram overlap, or approximate substring search) against the source document to find the closest matching passage.
  3. Attach citation to cell: Whatever passage scored as the best fuzzy match gets attached as "the citation" for that cell — but since matching happens on the whole generated text rather than per-sentence, the granularity is coarse

How can citations at a sentence level be generated?

This isn't explicitly outlined in the text some potential ways include:

  1. Prompting-based: The model could be instructed via system/prompt engineering to output structured citations (e.g., "for each sentence, output the corresponding source index") as part of its generation — essentially asking the model to self-report indices as a formatting/output requirement.
  2. Constrained decoding / structured output: Rather than relying on the model to "remember" to cite correctly via prompting alone, the system could constrain the model's output format at the token level (e.g., forcing it to emit index tokens at sentence boundaries), which is more reliable than prompting alone but requires deeper integration with the inference pipeline.
  3. Retrieval-augmented indexing: The source document could be pre-chunked and indexed (e.g., by sentence, with unique IDs), and the model's job is to select from a known, indexed candidate set of sentences rather than generate a citation from scratch — reducing hallucinated citations by construction.
  4. Hybrid: Prompting to elicit reasoning + a separate verification/alignment step that checks whether the model's claimed index actually matches the claim.

Sharing the reasoning with users seems kind of obvious. Is there any reason why they didn't do this in the previous iteration?

  1. Latency/cost?
  2. Risk of hallucination?

“This also resulted in a boost in output speed. By reassessing our execution pipeline, we were able to adopt an implementation that kept per-cell latency low and the overall architecture simple while achieving sentence-level granularity. At scale — a 30-column, 1000-document table produces 30,000 concurrent cells — the latency of this approach is substantial.“

I don't understand why this new approach is faster. Where is the latency boost coming from exactly?

They don’t explain it in the blogpost but it could be because:

  1. Parallelization: processing citation-lookup across sentences concurrently rather than sequentially per cell
  2. Reduced post-processing: eliminating the separate fuzzy-matching pass (which requires comparing generated text against source text as a distinct step) by having the model point to indices directly during generation — removing an entire post-hoc computation stage

Let's say I've built this pipeline and I want to get applied legal researchers' resources to assist in the evaluation. How can I, in advance of securing these resources, estimate how much manpower I need? One approach I had in mind was using synthetic queries in LLM as a judge to get a preliminary estimate of the improvement, and then use that to calculate the sample size I need for statistical significance at a desired margin of error.

This is probably reasonable but we can enhance this by estimating p̂ (preliminary win-rate) on a much smaller sample (30-50) which we can then use to calibrate the LLM judge by computing the agreement rate.

“Reliability measures the rate at which Harvey fails to provide the user with an answer for parsing reasons or because of unexpected model outputs.”

How do you measure “unexpected model outputs”?

Not explained in the blog post, but could feasibly include:

  1. Schema/format violations — the model was expected to return output in a structured format (e.g., answer + reasoning + citation index), but instead returned malformed structure, missing fields, or extra unparseable text.
  2. Invalid citation indices — the model points to a document index/offset that doesn't exist or is out of bounds, which would fail downstream parsing even if the "answer" text itself looks fine.
  3. Refusals or non-answers — model declines to respond, or returns something like an apology/meta-commentary instead of the expected output type.
  4. Truncated or incomplete generations — output cut off mid-structure (e.g., due to token limits) such that required fields never complete.
  5. Encoding/formatting anomalies — unexpected characters, broken JSON, or malformed markup if the pipeline expects a specific serialization.

Most of these can probably be measured/detected by LLM as a judge or some simple heuristics.

“We measured quality through side-by-side preference testing.”

Is there any concern about just comparing the two of them and asking which one is better, as opposed to quantifying the degree of improvement? If I want to justify the significant resources dedicated to this project, the current setup (which doesn’t distinguish between narrow wins and big wins) seems like it makes it a bit difficult.

We could try a graded preference scale: "A much better," "A slightly better," "Tie," "B slightly better," "B much better." This is still a comparative judgment (avoids the calibration problems of absolute scoring across raters) but recovers magnitude information that pure binary throws away.

How would I convince a business leader that this project was worth doing?

I think some way to get a sense of the time-savings from the improvements would be helpful (maybe ask ALRs to estimate).

If we’ve had previous user research showing that latency is a driver of churn, reduced latency can be presented as a major win. It’s fair to assume that improved response quality (as detected by the ALR preference data) is good for the bottom-line (increased retention, positive word-of-mouth etc.).

How would you improve this system if you were building it today?

  1. Citation verifications step (latency and cost trade-off though)
  2. More structured reasoning trace
  3. Selective reasoning depth depending on what kind of cell output is being considered
  4. Task-based outcome metrics in evaluation (actual time-to-complete a review task, error/override rates in production)
  5. Failure handling: if the primary pipeline produces an unparseable or low-confidence output, retry with a simpler/more constrained generation mode (e.g., force extraction-only, no reasoning) rather than surfacing a hard failure to the user
  6. Tiered latency budgets by query complexity. Not all cells need the same processing depth; routing simple yes/no extractions through a faster path and reserving the full reasoning+citation pipeline for interpretive questions would reduce average latency without sacrificing quality where it matters

Level of understanding: 🙂

2025 Blogpost

Building The Intent Engine: How Instacart is Revamping Query Understanding with LLMs

The historical solution faced various challenges: Struggled with broad search queries like “healthy food” or “frozen snacks” Lack of good labeled data - no direct feedback from clicks/conversion and pseudo-labels derived from user…

My notes

Summary

Context

Instacart users can search for grocery items on the platform. User searches are often unstructured and don’t map neatly to item names. This has historically been handled by various traditional machine learning models used for query classification, query rewrites, semantic role-labeling, etc.

Problem

The historical solution faced various challenges:

  1. Struggled with broad search queries like “healthy food” or “frozen snacks”
  2. Lack of good labeled data - no direct feedback from clicks/conversion and pseudo-labels derived from user behaviors are inherently noisy
  3. Struggled with highly specific or rare searches like “red hot chili pepper spice” or “2% reduced-fat ultra-pasteurized chocolate milk” because of data sparsity
  4. System Complexity - they maintained multiple independent models for individual QU tasks which introduced lots of overhead

Solution

They used LLMs to consolidate and enhance QU (query understanding) models

  • Query category classification: original model provides shortlist and LLM re-ranks with injected Instacart context
  • Query rewrites: three distinct rewrite types: Substitutes, Broader queries, and Synonyms. Each type is handled by a dedicated prompt with advanced prompt engineering
  • Semantic Role Labeling (SRL): they want to extract structured concepts from a user query, such as product, brand, and attributes. They ask a big LLM to generate cache for head queries and training data for a smaller LLM to handle the rest

Some key insights/new things I learned

  1. A more complex LLM can be used for caching head queries and generating synthetic data for teacher-student distillation simultaneously
  2. Query rewriting can be split into different types
  3. Post-LLM generation guardrails through semantic similarity filters can be helpful
  4. The right company-specific context injected into the prompt can improve performance a lot
  5. [Not sure if they actually do this but could be cool] for low-cardinality queries that have higher-cardinality embedding space neighbors, you could use the neighbors’ cached LLM results instead

Design Details

  • Data
    • User session and conversion history data, catalog data
    • Synthetic data for SRL from big LLM
  • Architecture
    • [Explained in Summary > Solution]
  • Evaluation
    • Metric(s): Precision, Recall
    • Data: User session and conversion history data
    • Results: Improvements observed but magnitudes relative to baselines not specified
  • Deployment
    • Optimizations
      • Adapter Merging & Hardware Upgrade
      • Quantization explored but rejected
      • GPU autoscaling
    • [Not described]
  • Post-deployment
    • [Not described]

What I didn’t understand or am unsure about

Question

Answer

How does query understanding fit into the broader search pipeline? What are the remaining stages?

Piecing together Instacart’s other engineering posts, the pipeline is roughly:

  1. Query understanding — normalization/spell correction, category classification, rewrites, SRL tagging
  2. Retrieval/recall — hybrid: lexical (Postgres full-text search, GIN indexes, a customized ts_rank), embedding-based retrieval (their ITEMS bi-encoder over a FAISS ANN service), and category-based retrieval
  3. Merge/fusion — reciprocal rank fusion or a weighted convex combination of lexical and semantic scores
  4. Ranking — multi-objective, balancing “relevance, popularity, and personal preference”; the embedding score is a feature here
  5. Serving layer — ads insertion, store-level availability filtering, dedup, blending, pagination
  6. Post-serving — out-of-stock replacement suggestions, feedback logging

What kinds of “traditional machine learning models” have historically been used for search-related tasks at Instacart for the different use-cases they outlined?

[Post] Only two are named, both in the Fig 1 caption: query classification “relied on a FastText model for multi-label classification,” and query rewrites “were generated by a separate system that mined user session behavior” — which is a mining/statistics pipeline, not really a model.

[Post] For category classification the legacy setup is described as a “massive multi-class classification problem” predicting “the top-K most likely categories from a flat list.”

[Search] Adjacent Instacart posts name ITEMS (Instacart Transformer-based Embedding Model for Search), a Sentence-Transformers bi-encoder fine-tuned on their conversion data, and an earlier MiniLM-L3-v2 bi-encoder used for hybrid retrieval. Both are retrieval-side, not QU.

[Inferred] SRL is never given a legacy model. The pre-LLM standard for e-commerce query tagging is BIO sequence labeling — CRF, BiLSTM-CRF, or BERT token classification. Reasonable guess, unstated in the post.

Why would the methods they've historically used struggle with broad queries?

  • No lexical bridge. FastText is a bag-of-n-grams linear classifier over the surface string. “healthy food” shares no characters with Yogurt, Produce, or Whole Grains. The model can only memorize the association from conversion data — it can’t derive it.
  • The label distribution is genuinely flat and multi-modal. Top-K with a fixed threshold either cuts real categories (low recall) or admits noise (low precision). There’s no single correct answer to fit toward.
  • It learns popularity, not semantics. Conversions for a broad query are dominated by whatever is merchandised or best-selling, so the model converges on “what people buy” rather than “what the query means.”
  • No compositional reasoning. “healthy” is a constraint over attributes (low sodium, high fiber, no added sugar), not a category. Deriving that requires world knowledge the model doesn’t have.

“QU operates upstream and doesn’t benefit from direct feedback like clicks or conversions. The pseudo-labels we derive from user behaviors are inherently noisy.”

Wouldn't this be a problem at any layer of the stack, regardless of what methods you use?

It’s a matter of degree rather than kind — but the degree is large enough to change what’s possible:

  • Ranking gets directly attributable labels. An item was shown at position 3 and clicked. The label sits on the same object the model scores.
  • QU’s outputs are never shown to a user. A conversion is credit-assigned through retrieval → ranking → UI. If someone searches “bread” and buys bananas, you can’t separate a bad rewrite from bad retrieval from bad ranking from a user who simply changed their mind. That’s credit assignment through a long, non-differentiable pipeline.
  • Confounding is worse upstream. QU’s own output determines the candidate set, so the feedback you observe is censored by the thing you’re trying to evaluate. Ranking has a version of this — position bias, selection bias from retrieval — which is why IPS and click models exist. But ranking has a label to debias; QU only has a proxy.
  • QU labels are set-valued and partly subjective. “What are the correct categories for ‘healthy food’?” has no unique answer, which is why the fallback is human evaluation — expensive, and the reason the post reaches for LLMs at all.

“This heterogeneity introduced inconsistencies”

What kinds of inconsistencies?

The realistic failure modes:

  • Semantic disagreement between components. The classifier maps “gluten free bread” to Bakery while the rewriter expands it to “sourdough” — which contains gluten. Two parts of QU assert contradictory things about the same query.
  • Taxonomy version skew. Each model trained against a different snapshot; categories get renamed, split, or merged, so one model emits IDs another considers dead.
  • Preprocessing skew. Different tokenization, casing, stemming, and spell correction per system, so “zip-lock” and “ziplock” are treated identically by one component and differently by another.
  • Training-window skew. Rewrites mined over 30 days, classifier trained over 90 — different implicit definitions of “popular.”
  • Metric incomparability. One team’s precision is human-labeled, another’s is a conversion proxy. You can’t reason about end-to-end quality because the numbers don’t compose.
  • Training/serving skew duplicated N times, each with its own bugs.

“slowed down development cycles”

Why?

  • N systems × (data pipeline + training + eval + serving + monitoring) is N times the maintenance and on-call burden, with no shared leverage.
  • A change to a shared input — a taxonomy revision, a new catalog field — forces N retrains and N separate validations.
  • Each system needs its own labeled eval set, and building one is the long pole. Every new task starts from zero.
  • No transfer. Improving the rewriter’s handling of “organic” does nothing for the classifier. Effort doesn’t compound.
  • Launching a new QU capability means a new bespoke model — a multi-quarter project — versus writing a prompt.

“Trained on diverse textual data, LLMs possess world knowledge that enables them to make logical inferences from user queries”

Isn't this also true for models like BERT, which I'm assuming were part of their historical pipeline?

Yes, BERT has world knowledge — but three differences are load-bearing:

  • Scale. BERT-base is 110 million parameters trained on roughly 3 billion words. A frontier LLM is three to four orders of magnitude larger on both axes. “Italian parsley ≈ flat parsley” is exactly the kind of long-tail fact that only survives at that scale.
  • Access mode — the decisive one. BERT’s knowledge lives in its representations, and the only way to extract it is a trained head on your labeled data. If you have no clean labels (Q4), you cannot get at it. An LLM exposes the same knowledge through natural-language instructions, requiring zero labels. That’s the whole unlock.
  • Generation and steerability. Rewrites are a generative task — BERT can tag spans but cannot produce “curly parsley” as a substitute; it has no decoder over that space. And instruction-following is what makes “generate a broader query, not a synonym” expressible at all. That single capability is what makes the three-rewrite-type design in section 2 possible.

Short version: BERT has world knowledge. It doesn’t have label-free, generative, steerable world knowledge.

“We build data pipelines that retrieve and inject Instacart-specific context”

What are good ways to determine what context would be most helpful for a given task?

A practical method, in order:

  1. Build an error taxonomy first. Sample 100–200 failures and label why each was wrong. Missing-knowledge errors (the model doesn’t know “verdant machine” is a brand) → context helps. Reasoning or format errors → context won’t help; prompting or fine-tuning will. This one step prevents most wasted effort.
  2. The human-oracle test. Could a smart new hire with no Instacart data access answer this? If not, what would they look up? That lookup is your retrieval source. Cheapest, highest-signal step available.
  3. Leave-one-out ablation. Build context as independent blocks — conversions, catalog, taxonomy, session — and ablate each against a fixed eval set. Keep only what moves the metric; every block costs tokens, latency, and adds distraction risk.
  4. Prefer high-precision, low-volume context. Top-5 converted categories beats top-50. Irrelevant context measurably degrades LLM output, so more is not better.
  5. Check for label leakage. If your context essentially contains the answer for head queries, your eval overstates gains on the tail — precisely where you need it to work.
  6. Make sure it works cold. The tail queries you care about have no conversion history by definition. Include at least one source that works with zero engagement data — which is exactly why Instacart pairs conversion history with catalog embedding similarity.

“For the most advanced use cases, we fine-tune models on proprietary data”

Is there a good way to know in advance or in a less resource-intensive way whether fine-tuning is likely to be helpful for your use case before you dedicate the full set of resources needed for it?

Cheapest-first sequence:

  1. The ceiling test. Can a frontier model with your best prompt and context solve the task well? If no, fine-tuning a smaller model almost certainly won’t — fine-tuning transfers behavior, it doesn’t invent capability. If yes, you have a teacher, and the problem reduces to a known-feasible engineering exercise. Instacart is exactly this case, which is why their story works.
  2. Name what fine-tuning is for. Three distinct goals with very different success rates: (a) cost/latency — distill a known-good big model; high success. (b) format and style adherence — very high success, small data. (c) new knowledge — low success; RAG is usually the better tool.
  3. Measure the few-shot slope. Run 0-shot → 5-shot → 30-shot. A steep improvement means the model learns this task from demonstrations, and fine-tuning is just many more demonstrations. A flat curve means the bottleneck is knowledge or capability, and more examples won’t fix it.
  4. Run a small LoRA pilot. 1,000–3,000 examples, a few GPU-hours. The curve from 500 → 1,000 → 3,000 tells you whether more data pays. Hours, not quarters.
  5. Check data feasibility before compute feasibility. Can you produce 5,000–10,000 high-quality examples without human labeling? If not, the project’s real cost is annotation, and that’s what usually kills it.
  6. Check task stability. If the label definition churns monthly (taxonomy revisions), fine-tuning bakes in something you’ll repeatedly have to re-bake. Prefer RAG.

“Our legacy approach treated this as a massive multi-class classification problem”

Is every element of every level of the taxonomy hierarchy a class in this model? Would something be able to classified as both “Meat” and “Short Ribs”?

[Post] Yes to both, and the post’s own example proves it: for “butter milk” the legacy model predicts (“Dairy”, 0.95) and (“Milk”, 0.92) — explicitly described as “distinct, non-hierarchical outputs.” So the label space is flat and spans levels, and an ancestor plus its descendant can both be returned, because the model has no representation of the hierarchy.

Terminology flag: the post is internally inconsistent. The body says “multi-class”; the Fig 1 caption says “a FastText model for multi-label classification.” The behavior described — top-K, multiple simultaneous labels — is multi-label. Multi-class means exactly one label.

[Inferred] The pathology this creates: independent labels mean probability mass splits across redundant categories, the top-K budget gets spent on duplicates, and you get no signal about how specific to be. The new approach fixes this by predicting a category path (“the LLM’s predicted category path”), which is hierarchy-aware by construction — one output encodes both the level and the lineage.

“It directly powers recall and ranking, helping us retrieve items from the right categories and intelligently expand the search when a query is broad or ambiguous”

  1. By “right categories”, do they mean they apply filtering?
  2. How does search expansion work?

(a)

[Search] Instacart’s embeddings post lists “category-based retrieval” as a retrieval source alongside keyword and embedding retrieval.

[Inferred] Almost certainly not a hard filter in the general case. Two likelier uses:

  • As a candidate generator — predicted categories trigger an additional retrieval pass (“fetch popular items in category C”), whose results merge with lexical and embedding candidates.
  • As a ranking feature — item’s category is in the predicted set → boost.

Hard filtering on a QU prediction is dangerous: one classification error zeroes out recall for the entire query. Filtering is plausible only for high-confidence narrow queries, or in combination with the similarity guardrail.

(b) [Inferred] Expansion here means widening the candidate pool across taxonomy branches, not rewriting the text. For “healthy food,” many categories score highly, so you retrieve from each and blend — which is also what produces category-shelf-style result pages for broad queries. Note this is a different mechanism from the query rewriting in section 2, even though both get called “expansion.”

“First, being trained on noisy conversion data (e.g., a user searches “bread” but buys bananas) means it can produce irrelevant suggestions”

Shouldn’t the noise average out and thus be mitigated?

Averaging only rescues you if the noise is zero-mean and independent of the label. Neither holds here:

  • The noise is systematic, not random. “bread → bananas” happens because users do multi-item shops in one session. The mislabeling correlates with what’s popular overall — bananas are among Instacart’s top items. So every query accumulates bias toward the same popular categories. More data makes that bias more confident, not smaller. This is the crux: you’re not averaging away noise, you’re converging on a biased estimate.
  • There’s a feedback loop. The current system determines what’s shown → what converts → tomorrow’s labels. The noise is self-reinforcing, not i.i.d.
  • Volume asymmetry. Averaging needs samples. Head queries have plenty, so it looks fine there. Tail queries — the entire target of this project — have a handful of conversions, so variance never washes out. This is exactly why the post pairs “noisy labels” with “data sparsity” as separate but compounding problems.
  • Precision is the operative metric. Averaging shrinks mean error but doesn’t remove the tail of visible failures. A small persistent probability mass on a wrong category still surfaces to a user as a visibly bad result.

“we retrieve the top-K converted categories for each query as initial candidates”

Wouldn’t this struggle with low cardinality queries? Do they do any kind of rewriting or normalization to mitigate that?

[Post] Not addressed. No normalization or backoff step is described on the classification path.

[Inferred] This is a genuine gap, and arguably the weakest seam in the section. If a query has three historical conversions, the candidate set is tiny and possibly wrong — and a re-ranker cannot recover a category that was never in the candidate list. The LLM’s world knowledge is wasted if the retrieval step already threw away the right answer.

Standard mitigations, all of which Instacart plausibly has but doesn’t mention:

  • Back off to a rewritten/normalized form. They already generate synonyms and broader queries — retrieve conversions for those instead. The post never connects the two systems, but it’s the obvious composition.
  • Back off to embedding neighbors. Use the query embedding to find similar historical queries and pool their converted categories. They do exactly this on the SRL path (“brand names with high semantic similarity, ranked by embedding scores”), so the machinery exists.
  • Back off to item-level signal. Retrieve items by text/embedding, then use those items’ categories as candidates. Works with zero query history.
  • Let the LLM generate rather than re-rank when candidates are thin, with the guardrail catching invalid paths.

Note the structural asymmetry: the classification section is head-query-shaped, and the SRL section is where they explicitly solve the tail. The post presents both as “the LLM approach” without flagging that only one handles cold start.

“we use an LLM to re-rank them with injected Instacart context”

  1. What context?
  2. Do they use structured outputs? If not, how to ensure ranking output format is correct?

(a) For category re-ranking you’d want: candidate category paths with human-readable names rather than IDs; conversion counts or rates per candidate; a few exemplar item titles per category (this is what disambiguates a category whose name is vague); and the taxonomy’s own description of each node.

(b) In rough order of robustness:

  1. Constrained decoding / grammar — vLLM guided decoding, Outlines, XGrammar. Invalid output is structurally impossible.
  2. Provider-side JSON schema mode if using a hosted API.
  3. Rank by index into the candidate list rather than emitting category strings. This makes hallucinating a nonexistent category impossible by construction and shrinks the output to a few tokens. For a re-ranking task specifically, this is the cleanest option and I’d expect it in a mature system.
  4. Plain JSON prompt + parse + retry — weakest, but workable offline.

One clarification worth making: their guardrail (semantic similarity filter) is not a format check. It validates relevance, not schema. A malformed response fails before it ever reaches the guardrail.

Also: this is an offline/batch path, so retries are cheap and the format-safety pressure is far lower here than on the real-time SRL path.

“finally, we apply a post-processing guardrail. This filter computes a semantic similarity score between the embeddings of the original query and the LLM’s predicted category path, discarding any pair that falls below our relevance threshold.”

What model do they use for this score? Wouldn’t it struggle for the same reasons their original approach did?

[Search/Inferred] Two obvious in-house candidates: ITEMS (their Sentence-Transformers bi-encoder fine-tuned on conversion data) or the older MiniLM-L3-v2. ITEMS is the likelier one — already trained on query↔item pairs.

[Inferred] On the objection — the guardrail mostly escapes the original failure modes, for three reasons:

  • Much easier task. The legacy classifier had to choose among thousands of categories. The guardrail only has to reject an obviously unrelated pair. A crude, low-resolution filter can be valuable where a crude classifier is not.
  • Decorrelated errors. The LLM fails by hallucinating plausible-sounding but irrelevant categories; those are usually lexically and semantically far from the query, which is exactly what an embedding catches. The failure modes don’t overlap, which is what makes the two-stage design work.
  • Tunable in the safe direction. Set the threshold loose so you only kill clear garbage. You’re not asking it to be right; you’re asking it to be right about egregious cases.

But the instinct is partly correct in one place: if the guardrail’s embedding model was trained on the same noisy conversion data, it inherits the same popularity bias — and may wrongly reject correct but rare query-category pairs. That would silently degrade exactly the tail queries this whole project exists to fix. The post doesn’t discuss this, and it’s the risk worth measuring before shipping.

“Our legacy system mined candidate rewrites from user session data”

What does this mean?

[Inferred] The standard technique is logging within-session query reformulation pairs. A user searches q₁ = “gluten free bread,” gets poor results, reformulates to q₂ = “gf bread,” and converts. You record (q₁ → q₂). Aggregate over millions of sessions, score by frequency and by the conversion lift of q₂ over q₁, then threshold.

Common variants:

  • Full reformulation chains within a session, not just adjacent pairs
  • Click/purchase-graph co-occurrence — two queries that lead to the same items are related
  • Abandon-then-succeed pairs — highest signal, because the user has effectively labeled q₁ as insufficient

[Inferred] Why it caps near 50%: it can only produce a rewrite for a query that (a) someone has typed before and (b) was followed by a reformulation. Novel and tail queries have neither, by definition. It’s a memorization system with zero generalization — which is precisely the gap an LLM fills.

“This proved too ambiguous.”

How would they assess whether it was too ambiguous?

[Inferred] That reads like a qualitative, eyeballed finding, not a measurement. The rigorous versions:

  • Type-distribution audit. Sample N queries, have raters label each generated rewrite’s type (synonym / substitute / broader / other) and usefulness. If output collapses onto synonyms — say 70% when you needed a spread — the instruction is under-specified. This directly diagnoses “ambiguous.”
  • Incremental recall — the cheap quantitative version. How many items does the rewrite retrieve that the original didn’t? “one percent milk” adds roughly zero. No humans needed, and it operationalizes “not useful” exactly.
  • Output-diversity check. Run the same prompt repeatedly at temperature; tight clustering means the model has silently picked one interpretation of an ambiguous instruction.

The generalizable lesson: “generate rewrites” doesn’t specify an objective. Splitting into three types is really specifying three separate objectives so each becomes individually evaluable — the evaluability is arguably the bigger win over the quality.

“This led us to design specialized prompts for three distinct rewrite types”

Are these three rewrite types all done every time for every query or is each of them triggered by its own unique conditions being met?

[Post] Not stated. It says each type is “handled by a dedicated prompt,” and that coverage rose above 95% “with 90%+ precision across all three types” — implying all three are generated broadly.

[Inferred] The likely design is generate all three offline, trigger conditionally at serve time:

  • Generation: run all three across the precomputed query set. You can’t know in advance which you’ll need, and offline compute is cheap. “Coverage over 95%” is a claim about the generated inventory, not about firing rate.
  • Serving: the post says rewrites matter “especially when the original query does not return sufficient results.” So the natural trigger is result-count or result-quality based. If the original returns enough good items, use nothing. If thin, escalate in order of intent-drift risk: synonyms first (safest, preserves intent) → broader (widens, some drift) → substitutes (changes the product, riskiest).

That escalation ladder is inference, not theirs — but it’s the only ordering that makes sense given the risk profile of each type.

“increased our query rewrite coverage to over 95%”

This feels like a very gameable metric in isolation?

Note they do pair it — [Post] “over 95% with 90%+ precision across all three types.”

[Inferred] Why coverage alone is nearly meaningless here: an LLM emits something for any input. Coverage measures the model’s willingness to answer, not quality. You can drive it to 100% by removing your abstain path. The legacy 50% was low because a mining system genuinely couldn’t produce candidates; an LLM’s pre-filter coverage is ~100%. So the honest framing is “50% from mining” vs “95% surviving guardrails” — and the interesting number is how much the guardrails removed, which they never report.

What would make it credible:

  • Coverage on the tail specifically. Head coverage was never the problem, and an aggregate number is dominated by head volume.
  • The coverage/precision curve as the guardrail threshold sweeps. A single point on that curve is unfalsifiable — you can always pick the flattering point.
  • Incremental recall and conversion lift on queries where a rewrite actually fired.
  • Null-result-rate reduction — the thing rewrites exist to fix, and the metric a business would actually pay for.

Coverage is an input metric. Nobody cares about it if it doesn’t move an output metric.

“90%+ precision across all three types”

How is this measured? Example?

Given the post’s own claim that QU lacks “direct feedback like clicks,” this is almost certainly human evaluation on a sampled set — possibly LLM-as-judge calibrated against human labels. The likely protocol, with the type-specific question a rater would answer:

  • Synonym — “Would a shopper consider these the same product request?”
    “1% milk” → “one percent milk” ✓ · → “skim milk” ✗ (different product)
  • Substitute — “If the original were unavailable, is this an acceptable alternative?”
    “Italian parsley” → “curly parsley” ✓ · → “cilantro” ✗ (different flavor profile)
  • Broader — “Is this a strict generalization that still contains the original intent?”
    “2% reduced-fat chocolate milk” → “chocolate milk” ✓ · → “beverages” ✗ (too broad) · → “whole milk” ✗ (sibling, not parent)

Precision = rewrites judged correct ÷ rewrites emitted, computed per type. Recall isn’t reportable because there’s no enumerable ground-truth set of all valid rewrites — which is exactly why they fall back to coverage as a recall proxy, and why Q19’s critique lands.

[Inferred] Important caveat: precision defined this way measures semantic correctness, not usefulness. “one percent milk” would pass a synonym-precision check while being precisely the failure they described one paragraph earlier. The metric they report does not capture the problem they say they solved.

“we are now adopting context engineering to make rewrites more convertible, personalized, and session-aware”

What would those improvements look like in practice?

[Post] One concrete mechanism: “injecting user engagement signals, such as the top-converting product categories from their subsequent searches in the same session.”

[Inferred] Unpacking each:

  • Convertible — bias rewrites toward products that actually sell and are actually in stock at the user’s selected retailer. Concretely: score candidate rewrites by the historical conversion rate of the items they retrieve, and gate on store-level availability. “Italian parsley → curly parsley” is worthless if that store stocks neither.
  • Personalized — condition on purchase history and dietary pattern. “milk” from a user who has only ever bought oat milk should rewrite toward plant-based, not dairy. Same for brand affinity: a repeat buyer of one brand gets rewrites that preserve it.
  • Session-aware — condition on queries already issued this session. Someone who searched “ground beef,” then “ricotta,” then “noodles” is making lasagna, so “noodles” should rewrite toward lasagna sheets, not ramen. This is the same idea as the post’s closing ambition of separating “lasagna ingredients” from “quick lasagna recipe.”

[Inferred] The cost, which the post doesn’t mention: all three push work from offline-and-cacheable to per-request, because the context is now user- and session-specific. That breaks the head-query cache — the exact mechanism that makes their economics work (Q22, Q29). That’s very likely why this is framed as future work rather than shipped.

“populating a cache for our most common “head” queries”

How to determine which queries? What threshold? Is it dynamic/time-windowed?

[Post] Never specified. Only: “the power-law nature of search traffic,” and — crucially, buried in the takeaways — “smart caching meant only 2% of queries needed real-time inference.”

[Inferred] That 2% is the answer in disguise: the cache is sized to cover 98% of query volume. The natural construction:

  • Selection. Rank distinct queries by frequency over a rolling window; take the prefix covering 98% of impressions. Under a power law that’s a small fraction of distinct strings covering the overwhelming majority of traffic.
  • Window. Should be rolling (30–90 days) and refreshed on a batch schedule, or it can’t track seasonality — “pumpkin spice” is head in October and tail in April.
  • Dynamic promotion. Log cache misses; any query crossing a miss-count threshold gets queued for the offline teacher pipeline and promoted on the next run. This is what keeps the real-time model’s load pinned near 2% as traffic shifts.
  • Key normalization. The cache key should be normalized text (lowercase, whitespace, punctuation, probably spell-corrected), or hit rate craters on trivial variants.
  • Invalidation is the hard part they don’t mention: entries must expire when the taxonomy, catalog, or prompt changes, or the cache silently serves stale tags indefinitely. This is usually the biggest operational liability in a design like this.

What model used for Teacher and why?

[Post] Never named. Only “a powerful offline process,” “much larger frontier model,” “massive LLMs,” and Fig 5’s unlabeled “baseline.”

[Inferred] For work published November 2025, the plausible set is a GPT-4o/4.1-class, Claude Sonnet-class, or Gemini Pro-class hosted model, or a large open model (Llama-3-70B/405B) run in-house. Choosing Llama-3-8B for the student mildly suggests comfort with that family, but a hosted frontier API is the more common teacher because batch pricing makes it affordable.

Why a big model as teacher: you need the teacher’s quality to be the ceiling you distill toward. You only pay for it on head queries, in batch, with no latency constraint — so it’s expensive per call but amortized across a cached result serving millions of impressions. That amortization is the entire economic argument for the architecture.

Is their approach equivalent to standard teacher-student distillation or something different? Why this fine-tuning approach vs another?

It is not classical knowledge distillation, and the difference is worth being precise about:

  • Classical KD (Hinton et al.) trains the student on the teacher’s full output distribution — soft logits, KL loss, temperature. Requires logit access and a shared tokenizer/vocabulary.
  • What Instacart did is sequence-level distillation — take the teacher’s final generated text as a hard label, train with ordinary cross-entropy. Colloquially everyone calls this “distillation,” but structurally it’s supervised fine-tuning on machine-generated labels.

Why that variant:

  • It works through an API. No logits needed, so the teacher can be a closed frontier model.
  • It’s tokenizer- and architecture-agnostic. Teacher and student need nothing in common.
  • The training code is just standard SFT. No custom loss, no distillation infrastructure.
  • The decisive one: the teacher’s advantage comes largely from injected RAG context. Hard-label SFT is how you bake that retrieved context into the student’s weights — so the student doesn’t need to run retrieval at inference. That is what buys the latency. The student isn’t just a smaller model; it’s a model that has internalized a retrieval step.

One refinement worth naming: [Post] the teacher’s outputs pass “a post-processing guardrail [that] validates the tags against our catalog” before becoming training data. So it’s filtered / rejection-sampled distillation, which meaningfully raises label quality over raw teacher output.

Why LoRA over full fine-tuning [Inferred]: far less memory, hours instead of days, cheap to sweep variants, trivial rollback (keep one base, swap adapters). The cost is a lower capacity ceiling — which didn’t bind here because the task is narrow.

How did they know in the first place that the Teacher LLM works well for this task? Who evaluates the teacher?

[Inferred] The standard answer, and near-certainly what happened:

  • A human-labeled golden set. A few hundred to a few thousand queries with expert-annotated SRL tags. Expensive — which is why it stays small — but you build it once and it anchors everything downstream. This is exactly the “costly and time-consuming human evaluation” the post cites in its challenges section.
  • Stratified by traffic band (head/torso/tail) and query type (brand-led, attribute-heavy, ambiguous). A teacher that’s excellent on head and weak on tail is useless here, since the tail is the whole point.
  • The catalog guardrail as an unsupervised check. Tags that fail catalog validation are wrong by construction — a cheap, large-scale lower bound on teacher error, no humans required.
  • Downstream proxy. Do teacher-produced tags improve offline retrieval/ranking metrics? Noisier, but scales.
  • The A/B test is the final arbiter, but it validates the system, not the teacher in isolation.

The structural problem worth sitting with: Fig 5 compares the student against the frontier model — i.e. the student is scored against the teacher, not against ground truth. If the teacher has a systematic blind spot, the student inherits it perfectly and the chart still shows 95.7% F1. “Matches the teacher” and “is correct” are different claims, and only a golden set separates them. They report the first and imply the second.

“For the real-time SRL model, we fine-tuned an open-source Llama-3–8B model using LoRA (Low-Rank Adaptation). The model was trained on the dataset from the offline “teacher” pipeline.”

What does this fine-tuning process look like concretely? How do I set up the necessary tools, datasets, infrastructure etc for this?

[Inferred] Concrete recipe as I’d run it today:

1. Data. JSONL of {messages: [system, user, assistant]} — user = the query (plus any context you want internalized), assistant = the JSON tag output. Target 5,000–50,000 examples. Hold out 5–10%, stratified by query frequency band so you can read tail performance separately. Dedupe aggressively; near-duplicate head queries will otherwise dominate the loss.

2. Framework, easiest → most controllable:

  • Axolotl — YAML config, batteries included, best default choice
  • Unsloth — fastest single-GPU (~2× speedup), ideal for 8B on one card
  • HuggingFace TRL SFTTrainer + PEFT — most transparent, best when you need to customize
  • LLaMA-Factory — good UI/CLI, broad model support

All wrap the same underlying PEFT LoRA implementation.

3. Hyperparameters, in order of impact:

  • r = 16 (8–32 is plenty for narrow extraction); lora_alpha = 2r
  • Target all attention and MLP projections (q, k, v, o, gate, up, down). Targeting only q,v is the old default and it underperforms — this is the single most common mistake.
  • LR 1e-4 to 2e-4, cosine decay, 2–3 epochs, bf16
  • Mask the loss to completion tokens only — don’t train on the prompt

4. Hardware. One 80GB A100/H100 fits 8B LoRA comfortably at bf16 with gradient checkpointing. On 24GB, use QLoRA (4-bit base via bitsandbytes) — quality is close for extraction tasks. Cloud: Lambda, RunPod, Modal, or SageMaker. A run at this scale is hours and tens of dollars, not a capital expense.

5. Eval, tracked per epoch. Three metrics: per-field F1 and exact-match on held-out, catalog-validation rate, and JSON parse rate. Parse rate should reach ~100% within one epoch — that itself confirms format learning is done and remaining error is semantic.

6. Merge and serve. peft’s merge_and_unload() → save as a plain HF model → serve with vLLM or TensorRT-LLM. Benchmark p50/p95/p99 at realistic concurrency, not batch-size-1 — that gap is where latency projects die.

7. Guarantee the schema in production. vLLM guided decoding (XGrammar backend) or Outlines, so parse failures become structurally impossible rather than merely rare.

“Merging the LoRA adapter weights directly into the base model”

What does this mean?

[Inferred] LoRA trains two small matrices per targeted weight — A (r×k) and B (d×r) — and inference computes h = Wx + BAx: the base projection plus a low-rank correction on a separate side path. Since BA has the same shape as W, you can precompute W' = W + BA once and store it. Inference becomes h = W'x — a single matmul, identical in cost to the unmodified base model.

Why it’s meaningfully faster:

  • Eliminates two extra small matmuls per targeted module, per layer, per generated token. That’s roughly 7 modules × 32 layers × every token.
  • Small matmuls are launch-overhead-bound, not FLOP-bound — so the cost is wildly disproportionate to the tiny parameter count. This is why 30% is plausible from a change that touches about half a percent of the weights.
  • Removes the branching side path, which unblocks fused kernels and CUDA graph capture in the serving stack.

What you give up: you can no longer hot-swap adapters or serve many fine-tunes from one base-model copy in memory. A real cost in a multi-tenant setup; irrelevant here, since they serve exactly one.

Numerically it’s exact — merging changes nothing about model outputs beyond float rounding. It’s a free win when you don’t need adapter swapping.

What is typical acceptable latency in search systems like theirs? Given that they’re only focused on the first part of the pipeline, what latency budget are they allocated?

[Post] Their numbers: ~700ms out-of-the-box on A100 → 300ms target, achieved via adapter merging + H100. FP8 quantization would have cut another 10% (rejected for a recall drop). End-to-end impact: “only a marginal latency increase.”

[Search/Inferred] Industry context:

  • E-commerce whole-page search typically budgets ~200–500ms p95 end-to-end; past ~1s, conversion loss is measurable. The canonical citations (Amazon’s “100ms costs 1% of sales,” Google’s 500ms/20%-traffic result) are old and over-quoted, but they’re why everyone budgets tightly.
  • Within that, retrieval is usually tens of ms — [Search] Instacart’s own embeddings post cites under 8ms for on-the-fly query embedding — ranking is tens of ms, and the remainder is network, ads, and rendering.

[Inferred] The key structural point: 300ms for a single QU component is enormous — larger than many teams’ entire backend budget. It is only affordable because of the cache. At a 98% hit rate, the average added latency is about 0.02 × 300ms ≈ 6ms, which is why “marginal” is honest rather than spin. The 300ms is a tail cost paid by 2% of queries — and those are cold-start queries where the alternative was bad results anyway, so the user’s willingness to wait is higher and the regret from getting it wrong is larger.

That trade — accept a large latency hit on a small, high-regret slice — is the actual transferable design insight here, more so than the 300ms number itself.

Is “average scroll depth” just the number of items users scrolled past before clicking?

[Post] Defined only parenthetically: “(users find items faster).”

[Inferred] Roughly yes, but the exact definition changes the interpretation, and they don’t give it. The three common variants:

  • Rank of first engagement — position of the first item clicked or added to cart. Lower is better. Most likely what’s meant, given “find items faster.”
  • Max scroll depth — deepest item position rendered during the session, clicks or not. Measures effort, but conflates “found it fast” with “gave up fast.”
  • Impressions before first add-to-cart — counts items seen rather than rank; handles grid layouts and lazy loading more cleanly.

Why the ambiguity matters: −6% is clearly good under the first reading. Under the second it’s ambiguous — abandonment also reduces scroll depth. A user who sees garbage and leaves immediately scrolls very little. Ruling that out requires pairing it with a conversion or engagement-rate metric. The complaints figure partially covers it, but not cleanly.

[Inferred] A grocery-specific wrinkle: users often scroll deliberately to browse alternatives even when results are excellent (“what yogurts do you carry”). That makes scroll depth a noisier relevance proxy in grocery than in, say, a support search box — worth knowing if you ever borrow the metric.

How would I convince a business leader that this project was worth doing?

Cost

  • Reduced overhead from mode consolidation?

Benefit

  • More conversions -> more revenue

How would you improve this system if you were building it today?

[Inferred] throughout, roughly ordered by expected value:

1. Build the golden eval set first. The biggest gap in the entire writeup (Q26). Without human-labeled ground truth, the teacher is unvalidated and the student is only ever measured against it. A few thousand stratified, expert-labeled queries is the highest-leverage artifact in the system — it makes every downstream decision cheap and every claim falsifiable.

2. Unify the three QU tasks into one model. Their own takeaway is “consolidate, don’t complicate,” yet they still ship three separate paths with different architectures. One fine-tuned model emitting categories + tags + rewrites in a single structured response is one forward pass instead of several, and it eliminates the cross-component inconsistency they named as the original problem (Q5). Multi-task training on a shared backbone usually helps each task, since they draw on the same underlying understanding.

3. Close the loop from serving back to training. Log cache misses, guardrail rejections, and null-result queries; feed them to the teacher on a schedule; retrain the student periodically. Prioritize queries where student and teacher disagree — that’s where the information density is. Their architecture is a one-way pipeline; making it a cycle is cheap and compounds.

4. Replace the exact-match cache with a semantic cache. String matching misses “gluten free bread” / “gluten-free bread” / “glutenfree bread.” Embed the query, ANN-lookup against cached keys, accept above a conservative threshold. They already run an ANN service. Higher hit rate translates directly into lower GPU spend.

5. Test whether you need 8B at all. A 1–3B model (Llama-3.2-3B, Qwen-3-4B class) may match on a task this narrow, at a fraction of the latency and cost. One experiment either saves real money or justifies the 8B — both outcomes are worth having.

6. Add abstention / confidence. The student always answers. Fig 5’s 96.4% precision means roughly 1 in 28 outputs is wrong with no way to identify which. A model that can say “unsure” lets you route those to a fallback — broader retrieval, or a synchronous call to a larger model with a longer budget — instead of confidently mis-tagging.

7. Audit the guardrail for popularity bias (Q15). Check whether the similarity filter disproportionately rejects correct-but-rare pairs. If so, make the threshold frequency-adaptive, or replace pure embedding similarity with catalog-grounded validation — “does any real item match this category path?” — which they already do on the SRL side.

8. Instrument the business metric. Conversion and null-result rate on the affected slice, in the A/B readout.

9. Modernize serving. Built today: server-side guided decoding, speculative decoding (SRL outputs are short and highly patterned, so draft-model acceptance rates should be excellent), and continuous batching in vLLM would likely land under 300ms without the H100 upgrade — which changes the hardware budget conversation entirely.

10. On their stated direction — session-aware, multi-intent — it’s right, but it breaks the cache (Q21). I’d resolve that by splitting rather than merging: keep query-level tags cached and stateless, and add a separate, small, cheap session-intent model whose output is combined at serve time. Putting session state inside the cached artifact destroys the 98% hit rate that makes the economics work.

2019 Blogpost original

Lessons from building AI to Stop Cyberattacks

My notes

What I didn’t understand or am unsure about

Question

Answer

“Attackers may launch attacks from compromised accounts making them even harder to identify from regular email”

What signals can you actually use if this happens? IP address, communication style, etc?

Per-sender behavior — writing-style deviation against that sender's own history, send-time-of-day shift, sudden send bursts, recipients never previously contacted

Mailbox configuration — new forwarding rules; rules that auto-delete or bury replies in Archive/RSS (the classic tell — the attacker hides the victim's incoming replies); delegate and permission changes.

Authentication / session — sign-in geography and ASN, impossible travel, new device fingerprint, legacy-protocol auth (IMAP/POP, which bypasses MFA), OAuth app consents

The IP address situation is a bit more complicated. It’ll be the same as the regular account holder (and thus not usable) if the attacker uses webmail or the mobile app. It’ll be different (and thus usable) if the attacker uses stolen credentials with an SMTP client which is much rarer.

“are identity of parties in the communication often targeted?”

What does this mean?

The intended reading: is this person's role one that attackers commonly go after? It's a prior on the identity, not a property of the message.

Role targeting — CEO/CFO (impersonated in BEC), AP clerks (invoice fraud), HR/payroll (W-2 and direct-deposit diversion), IT helpdesk (credential resets), executive assistants (authority by proxy)

“1 in 100,000 emails is advanced spear-phishing“

How was this estimate generated?

How a number like this necessarily gets produced: numerator = attacks their own engine caught, plus customer-reported and analyst-confirmed cases, plus retro-hunts. Denominator = total messages processed. So it's an observed detection rate, not true prevalence — a lower bound, because false negatives are by construction uncounted.

Three things that move it a lot:

  • Denominator choice. Abnormal sits behind the existing gateway (API, post-delivery), so its denominator is already spam-filtered mail. That inflates the rate relative to raw internet email — the 65-in-100 spam is already gone.
  • Definition. "Advanced spear-phishing" is their category, no external standard.
  • Customer mix. Their customers are enterprises that bought advanced email security — plausibly more targeted than average.

“Account sign-ins, mail filters, and other account activity”

Do we typically have access to all of this information?

Yes, for the API-based model. Microsoft 365 exposes it through Graph API plus the Office 365 Management Activity API: Entra ID sign-in logs, mailbox audit logs, inbox-rule create/modify events, message content and headers, mailbox permissions and delegates, OAuth consents. Google Workspace has equivalents (Admin SDK Reports API for login/user-account audit, Gmail API).

Two caveats that matter:

  • Licensing and consent. Full sign-in log retention depends on the customer's Entra ID tier, and you only get the scopes the admin consented to at install.
  • Lag. Audit events surface with delay — historically minutes for the Management Activity API. So they are not available at the sub-second scoring moment for the message in hand. They feed the account-takeover detector and the feature store, not the inline decision. This connects directly to Q17.

“Malware in attachments”

How is this identified?

Four methods, cheapest to most expensive:

  1. Hash / reputation — SHA-256 against known-bad corpora (VirusTotal, internal). Milliseconds; catches only knowns.
  2. Static analysis — OLE/VBA macro extraction from Office files; PDF JavaScript, embedded files and launch actions; recursive archive unpacking; YARA rules; ML classifiers over PE headers and section entropy. Catches variants; defeated by packing.
  3. Dynamic detonation — run the file in an instrumented sandbox (CAPE/Cuckoo lineage; commercially Joe Sandbox, VMRay) and watch process spawns, registry writes, network callbacks. Catches unknowns; takes 30 seconds to several minutes, and evasion-aware malware sleeps or checks for VM artifacts.
  4. Multi-engine AV — several commercial engines in parallel.

“Adversarial attackers — To make matters worse, attackers actively manipulate the data to make it hard on ML models:”

For each of the methods listed here, what is one counter strategy each?

Figure from the post

“The precision must be very high — to build a product to prevent email attacks we must avoid false positives and disruption of legitimate business communications”

Can we have different thresholds for different remediation severities (quarantine, block, etc)?

Yes, and this is standard practice: decouple score from action. The model emits a calibrated probability; a separate policy layer maps score bands to actions. Thresholds should be monotonically increasing in action severity, because each action has a different FP cost:

Figure from the post

Three second-order points:

  • Vary by attack class too. A suspected payroll-diversion BEC against the CFO justifies a lower bar than a generic promo-shaped phish, because the FN cost differs by orders of magnitude.
  • Vary by recipient. Lower thresholds for finance, AP, and executives.
  • The post-delivery architecture makes this much cheaper. Because Abnormal acts after delivery via API, it can decide late and revise — retracting a message when link detonation finishes or when the same campaign gets confirmed at another tenant. An inline gateway must decide once, immediately, with whatever it has.

One dependency: bands are only meaningful if the score is calibrated to an actual probability

“Is this pattern of communication unusual in any way?”

What are a couple of examples of stuff that would get flagged here?

Six kinds, organized by what's anomalous about the graph edge:

  1. Edge never existed — first-time sender to this recipient, especially one that immediately asks for an action.
  2. Edge exists, wrong direction or role — this vendor has always emailed AP, now it's emailing the CEO. Or the CEO mails a junior employee three orgs away they've never contacted.
  3. Edge exists, properties changed — same display name, different sending domain than the previous 400 messages from that vendor. Or same person, new reply-to.
  4. Volume/timing — an internal account suddenly mails 200 external recipients; a user sends at 3am local when they never have.
  5. Missing expected structure — a "reply" that belongs to no thread the recipient has; a vendor invoice arriving off its normal monthly cadence.
  6. Tenant-level shape — a domain that has never touched this organization now contacting five finance employees at once. That's campaign shape, invisible at the single-message level.

“Our ensemble of detection models”

Why ensemble as opposed to single model?

Five reasons, roughly in order of how much they matter here:

  1. Adversarial robustness. A single model is one thing to reverse-engineer. The post says attackers A/B-test against filters — if they find the perturbation that flips your one model, they win everywhere. With detectors keyed on different data (text vs. graph vs. identity vs. links), evading all of them at once is much harder, and the evasion that beats one usually trips another.
  2. Heterogeneous sub-problems. Invoice fraud, credential phishing and lateral ATO phishing have different base rates, signals and label volumes. A single model trained on all of it is dominated by the common class and underfits the rare, expensive one.
  3. Precision-budget arithmetic. At a 1-in-a-million FP target, you can OR together several high-precision, low-recall detectors and accumulate recall without any one of them being perfect. A single model has to hit that operating point alone.
  4. Operational independence. Ship and roll back separately, attribute failures to a component, patch a novel attack in hours with a signature while retraining runs for weeks.
  5. Variance reduction. The textbook decorrelated-errors argument — the least important reason in this setting.

What do they mean by “redundant detectors”? Why do they use it and what does that look like in practice?

Redundancy ≠ ensembling. Ensembling combines weak signals into a better score. Redundancy means deliberately overlapping coverage: more than one independent detector capable of catching the same attack on its own. It's defense in depth for the classes where a miss is catastrophic.

Why: if detector A fails — new evasion, model drift, a feature-pipeline bug, a bad deploy — detector B still fires. The failure modes are meant to be uncorrelated.

For invoice fraud specifically, the stack would look like:

  • The multi-modal deep model scoring the message end-to-end
  • A dedicated invoice sub-model over attachment OCR and invoice layout
  • A stateful check, not a model: does the bank account on this invoice match what this vendor has historically been paid to?
  • Analyst heuristics — "reply-to ≠ from AND first-contact sender AND banking keywords AND urgency"
  • Signatures on known campaign infrastructure

“for the convenience of including multiple modalities of data inside the same model”

A GBDT could also take in text and image embeddings as features though, right?

Yes — mechanically nothing stops you. The post's stated reason is convenience, and it undersells the real difference.

Four differences, and the first is the substantive one:

  1. Gradient flow. Feed a GBDT frozen embeddings and the text encoder never learns from your detection labels. The figure shows the NLP sub-model's "Text Embedding" feeding a joint Neural Net — gradients propagate back into the encoder, so it learns representations tuned to this task (phishing register, urgency, payment language) rather than generic semantics. Trees cannot backprop into an encoder. This is the whole ballgame.
  2. Dimensionality. A 768-dim embedding is 768 axis-aligned split candidates to a tree. Embedding information lives in directions, not coordinates, and trees split one coordinate at a time. GBDTs are weak exactly in the dense high-dimensional continuous regime. Note the converse: on genuinely tabular features GBDTs usually still beat NNs — which is why the figure retains a separate Tabular Sub-Model.
  3. Variable-length data. An email has 0–N links, 0–N attachments, 0–N images. A tree needs a fixed-width row, so you hand-pool ("max link risk," "mean image score") and throw information away. A network can pool or attend in a learned way.
  4. Engineering — the post's stated reason. One model, one training loop, one serving artifact, instead of a pipeline of separate embedders whose versions must be pinned and re-materialized every time they change. That last part is precisely the pain in his other Abnormal post on re-scoring past attacks.

Honest counterpoint: GBDT + frozen embeddings is a genuinely strong baseline, is what the post recommends starting with, and is cheaper to serve. The upgrade only earns its keep once you have enough positive labels to fine-tune — and labels are scarce here because of the base rate.

“we must be very careful that the model performs well autonomously and can handle distributional shifts”

How do they ensure this happens?

Pre-deployment

  • Calibrate the score to a real probability (Platt/isotonic), correcting for negative subsampling
  • Backtest on held-out future time slices, never random splits. Random splits leak badly here because campaigns repeat within a window.
  • Re-score historical attacks with today's pipeline to confirm the new model still catches old ones.

At deployment

  • Shadow mode: score live traffic, take no action, compare against production.
  • Gradual rollout by tenant with a pre-agreed abort criterion.

Monitoring

  • Watch the inputs, not just outputs: per-feature distribution drift (PSI/KL), null-rate spikes, cardinality changes. Feature-pipeline bugs look identical to genuine drift and are far more common.
  • Watch firing rate per tenant at a fixed threshold — the earliest available warning.
  • Asymmetry to plan around: precision is observable within days (customers report false positives); recall is not observable at all. You need proxies — analyst-labeled samples, purple-team and simulated attacks, cross-tenant confirmation.

Structural mitigations

  • Per-tenant normalization. Features expressed relative to an org's own baseline transfer across orgs far better than absolute features.
  • Keep the heuristic/signature layer, which can be updated in hours when the model degrades.

“Study the attacks, trends, and false negatives”

What are the sources for false negatives?

Seven channels, by how you learn about the miss:

  1. End-user reports — the "report phishing" button in Outlook/Gmail. Highest volume, lowest quality.
  2. Customer SOC escalation — the security team finds it in their own hunt and forwards it.
  3. Incident-driven — post-breach investigation after money moved or credentials were used. Lowest volume, highest cost, most informative.
  4. Retro-hunting on new intelligence — you learn a domain or hash is malicious later and sweep historical mail. Includes cross-tenant: an attack caught at customer B lets you re-search A and C. Abnormal's later material calls this "federated intelligence."
  5. Delayed signal — the URL was benign at delivery and weaponized afterward (benign-then-swap). Only re-crawling links post-delivery surfaces these.
  6. Threat-intel feeds and industry sharing — ISACs, vendor feeds, public reporting.
  7. Phishing simulation vendors (KnowBe4, Proofpoint Security Awareness) — attacks in your stream with known ground truth.

The bias to name explicitly: every channel is biased toward attacks that were noticed. The FNs you never learn about are the successful ones never attributed to email at all. Measured recall is therefore optimistic by an unknown and unmeasurable amount — which is the deeper reason the post's base-rate figures in Q3 should be read as lower bounds.

“Build a portfolio of detection models, heuristics, and signatures”

What does signature mean here?

A signature is an exact-match rule keyed to a specific known artifact. It fires on identity, not on learned pattern — the narrowest, most brittle, most precise layer of the stack.

In email that means: a SHA-256 of a known-bad attachment; an exact URL, domain, or IP of known phishing infrastructure; a YARA rule matching a byte pattern inside a file; a specific sender address or subject string from an observed campaign; a structural fingerprint of a campaign's HTML template.

Why an ML-first company keeps signatures at all: response time and explainability. When a campaign is live right now, you push a signature in minutes; retraining takes weeks. And when a customer asks "why was this blocked?", a signature has a one-line answer.

“Ensure each sub-model is representing its sub-problem”

What are examples of sub-models representing sub-problems in this context?

Figure from the post

All four feed a joint Neural Net → "Is Attack?"

[I] The design point that's easy to miss: the sub-problem is the modality, not the attack type. Short text is split from long text because a display name and an email body need different tokenization and different encoders — a subject line has no sentence structure to model, and a 24-character display name in a BERT-shaped encoder is mostly padding.

[S] The lesson in full: "Ensure each sub-model is representing its sub-problem well before trying to combine them into a larger network, or you will never make progress." Meaning: validate each branch in isolation — can the Link Sub-Model alone separate phishing URLs from benign ones? — before wiring it into the joint net. Otherwise a broken branch is invisible inside a joint loss, and you can't tell whether poor performance is an architecture problem or one garbage input.

[I] Sub-problems you'd expect beyond the figure: an OCR/vision model over attachment and landing-page images, an invoice-layout model, a per-sender writing-style model, and a communication-graph embedding.

“combine results of all models for a final decision on whether the email is malicious”

What would this look like in practice?

Four combination mechanisms, which are not mutually exclusive:

  1. In-model fusion — the figure's approach. Sub-models feed a joint neural net; combination is learned, inside the network, and there's only one score at the end.
  2. Stacking — each detector emits a score; a small meta-model (logistic regression or shallow GBDT) takes those scores plus context (tenant, attack class) and outputs the final probability. Needs clean held-out data or it leaks.
  3. OR-of-thresholds — each high-precision detector has its own threshold; malicious if any fires. This is what "redundant detectors" implies, and it's the natural fit for recall-through-redundancy. Simple and explainable, but FP rates add.
  4. Policy layer on top — score bands map to actions (Q7), plus overrides: allowlists, VIP-recipient escalation, tenant policy, and "if signature X fired, act regardless of score."

Realistically it's a hybrid: a learned score from the core multi-modal model, OR'd with a small set of high-precision specialist detectors and signatures, then a policy layer choosing the action.

One design constraint worth stating: attribution matters operationally. Customers ask why a message was actioned, and SOC analysts need to triage. So the combiner has to carry the firing reason forward — which argues against fully opaque fusion as the only layer.

“All this must be done at latencies of less than a second”

Where did this latency constraint come from?

Abnormal isn't inline. Their own architecture page describes a "pure API architecture that operates post-delivery," "outside the mail flow," with "no MX record changes," and notes that messages "briefly exist in the mailstore before remediation — typically measured in seconds."

[I] So the real constraint is time-to-remediate, not protocol timeout. The message is already sitting in the user's inbox. The window between delivery and retraction is the window in which someone can click, and mobile push notifications make that window very short — users read on phones within seconds of arrival. Every second is exposure.

The tension this creates, which the post doesn't acknowledge: several things on its own feature list cannot fit in a second —

  • Sandbox detonation of an attachment: 30s to minutes (Q5)
  • Headless rendering of a JS landing page behind a CAPTCHA: seconds
  • Audit-log events (sign-ins, rule changes): minutes of lag (Q4)

So in practice this has to be a two-tier design: a fast path scoring on what's already available (text, headers, cached feature store, graph lookups) and a slow path that re-scores when enrichment lands — with the ability to retract a message after the fact. The post-delivery architecture is exactly what makes that second bite possible, and it's the same property that makes tiered thresholds cheap in Q7.

Read "less than a second" as an engineering target for the online scoring path, not an end-to-end SLA.

“How can we detect invoice fraud better? To do so we must deeply understand the natural language and images in an invoice to identify Abnormal patterns.”

What is invoice fraud exactly and why is it a challenge compared to the stuff they’re already doing?

What it is: getting a business to pay a real-looking invoice into an attacker-controlled bank account. Three variants:

  • Vendor email compromise (VEC) — the attacker compromises the vendor's mailbox, sits inside a genuine payment thread, and sends a real invoice with changed banking details. This is what the post's "sit on a compromised account for months... inserting an illegitimate invoice into a conversation about payment at just the right time" bullet describes.
  • Vendor impersonation — a lookalike domain of a real supplier; no compromise needed.
  • Fake vendor — an invoice from a supplier that doesn't exist, aimed at a large AP department paying on autopilot.

Why it's harder than what they already do — seven reasons:

  1. No malicious artifact. No malware, no credential-harvesting URL, no payload. Every signal a legacy gateway is built on is simply absent.
  2. The text is legitimate business text. "Please find attached invoice #4417, net 30." Body-level NLP has nothing to grip — which is precisely why the post says you must understand the invoice, not the email.
  3. The signal lives in the image/PDF. Detection requires OCR plus layout understanding: extract vendor name, remittance account/IBAN, amount, invoice number — a document-understanding problem, a different and much harder extraction task than reading an email body. That's what "images in an invoice" is pointing at.
  4. It requires cross-message state. You cannot classify the message in isolation. You need a per-vendor payment-detail store: "this supplier has been paid to account ending 8812 for two years; this invoice says 4409." That's a data-engineering problem, not a model problem.
  5. ~100× worse base rate. <1 in 10 million vs. 1 in 100,000. Far fewer positive labels, unchanged precision bar.
  6. Every authenticity check passes. In VEC the mail genuinely comes from the vendor's real domain, valid SPF/DKIM/DMARC, inside a real thread with real history. The graph features that catch first-contact phishing all report "safe" — the exact opposite of the Q8 signals.
  7. Highest per-incident cost, in both directions. Six- and seven-figure wire losses on a miss; blocking a real invoice stops the business paying its suppliers. Both FN and FP costs peak here simultaneously.

“How can we better identify account takeovers? This is a difficult anomaly detection problem”

Do they actually use traditional anomaly detection algorithms as either models or features?

I don’t think they use traditional anomaly detection algorithms.

  • The dominant industry pattern is per-entity baselining and deviation scoring, not off-the-shelf unsupervised detectors. You model each user's own history and score how far today deviates — z-scores, quantiles, or a learned per-user profile. That is anomaly detection in the statistical sense, but it isn't "run an isolation forest on it."
  • Why the classical algorithms underperform here: they find rare, and in enterprise email rare ≠ malicious. A salesperson traveling, a new laptop, a quarter-end send burst are all rare and all benign. At a one-in-a-million FP budget, an unsupervised detector's alert volume is unusable standalone.
  • So the realistic role is as features into a supervised model, not as the decision-maker. Anomaly scores ("how unusual is this login for this user") become inputs alongside signals trained on labeled attacks.

How would you improve this system if you were building it today?

The attack side changed more than the defense side.

  • LLM-written lures destroyed the fluency signal that 2019 NLP leaned on — no more grammar errors, and personalized spear-phishing now runs at spam economics.
  • The attack went multi-channel: the callback step moved to cloned voice, SMS, Teams/Slack, and calendar invites. Email-only telemetry now sees a fragment.
  • Targeting is automated: retrieval over public data to rank who has payment authority.

Seven changes, most to least important:

  1. Model intent, not content. Since fluency no longer discriminates, stop asking "does this text look phishy" and model the request: what action, by whom, of whom, with what authority, at what point in a thread. "First-contact sender asks AP to change remittance details" is attack-shaped no matter how well written. This is an extraction-and-reasoning task an LLM does well and 2019 embeddings could not.
  2. Replace the OCR→NLP pipeline with a vision-language model reading the rendered email, the attached invoice, and a screenshot of the landing page in one pass. That single component attacks all three of the post's stated open problems — invoice fraud, phishing sites, image-encoded text — at once.
  3. Learned graph representations. 2019 used graph features; today use a GNN or temporal-graph model over the tenant's communication graph so relationship anomalies are learned rather than hand-engineered. Abnormal's 2026 material mentions graph neural networks, so this has happened.
  4. Cross-channel identity fusion. One identity risk score across email, Slack/Teams, Zoom, calendar and the IdP — so a suspicious Teams message following a suspicious email is one incident, not two independent low-confidence events that each fall below threshold.
  5. Adversarial evaluation as a first-class pipeline. Continuously generate LLM-based evasions of your own detector and train against them; treat detector robustness as a monitored metric, not a periodic red-team exercise.
  6. Prompt-injection hardening. Any LLM component now reads attacker-controlled text. Email content must be data, never instructions — this is a new attack surface the 2019 design didn't have.
  7. Fix some of it with workflow, not classifiers. For the highest-cost class, out-of-band verification of remittance-detail changes against a vendor record removes more residual risk than any marginal model improvement. Worth saying plainly rather than treating everything as a detection problem.

What I'd keep unchanged — these aged well: start simple and add complexity only when you can show it's better; heuristics as model features; redundant detectors; per-tenant baselining; calibration and thresholding discipline. Modest GBDT-on-tabular too — trees still beat NNs on genuinely tabular features.

2024 Blogpost original

Building DoorDash’s product knowledge graph with large language models

When a merchant comes onboard at DoorDash, their internal SKU data is added to DoorDash’s retail catalog. This additional data needs to be standardized and enriched to ensure quality and compatibility with the existing data. Historically…

My notes

Summary

Context

DoorDash maintains a dataset of all products sold by new verticals merchants - merchants operating a business other than a restaurant, such as a grocery, a convenience store, or a liquor store. Within the retail catalog, each SKU, or stock keeping unit, is represented by a list of product attributes.

This dataset has downstream applications for search and personalization.

Problem

When a merchant comes onboard at DoorDash, their internal SKU data is added to DoorDash’s retail catalog. This additional data needs to be standardized and enriched to ensure quality and compatibility with the existing data. Historically, this process has been handled manually by human contractors; however, this is slow, expensive and low-quality.

Solution

Attribute extraction using an LLM.

Design Details

  • Evaluation
    • Metric:
  • Post-deployment
    • N/A

What I didn’t understand or am unsure about

Question

Answer

It’s a little confusing to me that parents of items can be both brands (Coke) and sub-brands (Coke Zero). Why isn’t regular Coke a sub-brand of Coke the way Coke Zero and Cherry Coke are?

One way to think about this is that in their knowledge graph, a Sub-brand is created to capture a deviation from the core product (e.g., "Zero" for no sugar, "Cherry" for flavor). Because "Regular Coke" is the original standard, it has no deviating attributes to define it against the parent. It is the definition of the parent brand itself.

Depending on how the knowledge graph is used, having a Coke -> Regular Coke edge could increase latency without any actual benefit.

“Unstructured product description is passed to our in-house brand classifier”

What is their in-house brand classifier?

Exact model unclear from the blogpost but a reasonable assumption would be some sort of lighter-weight language model like BERT used as a closed-set classifier (trained on a limited set of labels).

“SKUs that cannot be tagged confidently to one of the existing brands are passed to an LLM for brand extraction.”

Let’s say they used BERT as their initial classifier. Does lack of confidence here imply a relatively flat distribution (high entropy) over output probabilities?

That’s one way to measure it. We could also use margin (difference between top two predicted probabilities).

“The extraction output is passed to a second LLM, which retrieves similar brands and example item names from an internal knowledge graph to decide whether the extracted brand is a duplicate entity.”

Why do we need a second LLM for this? And how is it retrieving the similar brands and item names? Is it through embedding similarity search for the extraction output or by navigating the knowledge graph in some way?

By second LLM, I think they just mean the same model but called a second time with a different prompt.

  1. The first call extracts the brand.
  2. The second call, augmented through the similarity search, checks whether the extracted brand already exists.

I think the retrieval is just through embedding similarity based on the RAG system they describe later.

“the in-house classifier is retrained with the new annotations”

Does it make sense to retrain every time a new brand is added? And what would the retraining look like exactly?

I think retraining every time a new brand is added is probably not computationally efficient.

For the retraining step, if the classifier is a fixed-output model (e.g., a Softmax layer with N brands), the "head" of the model would need to be expanded to include a new node for the new brand.

For language models, you could freeze the earlier layers and simply retrain the final few layers of the model on the additional data + some older examples (to avoid catastrophic forgetting). LoRA could also work well.

If the cost of full retraining seems really low though, might as well do that.

“Last year, we stood up a model to label all organic grocery products”

Does it make sense to have a separate model to do this labeling as opposed to having it be just Organic/Not Organic be one part of the taxonomy mapping of a larger model?

One of the main downsides to having a single model do all sorts of taxonomy mapping is certain features may exist for some items but not for other items. Food items can be organic/not organic but batteries can’t. Having a model try to figure out what the feature value is for a feature that doesn’t exist for the item being considered can increase computational cost and negatively impact performance.

{Worth thinking through the other tradeoffs and how to make a final decision}

Where are the training data labels coming from in the organic product labeling use-case?

Since they mention “highest precision”, they must have some labels.

Unclear but either human labels or powerful LLM (that is really good at generating labels but too slow/expensive to run in production).

“missed cases where organic is misspelled / dropped or has a slightly different presentation in the data”

What do they mean by a ‘slightly different presentation in the data’? That it’s not included in the product title but might be present in an image or some other metadata?

Could also mean something like “Org," or "Org.” or the word is concatenated with other text e.g., OrganicMilk or Milk(Organic).

I feel like the LLM reasoning approach could have been used for the entirety of their catalog. Using a distilled language model for what I’m guessing is < 1 million unique items doesn’t seem like it should be that slow or costly, at least not anymore. Is that fair?

# of unique items isn’t the right number to think about; unique merchant SKU data is. Different merchants can encode their data differently and each of these different encodings would have to be processed.

Determining what would actually make the most sense in production would need more information about the cost and latency structure.

“LLMs conduct online searches of product information and pipe the search results to another LLM for reasoning”

What exactly do they mean by ‘LLMs conduct online searches’ and how is the first LLM here different from the second one?

It could look something like this:

  1. The LLM is given a SKU (e.g., "Horizon Whole Milk 64oz").
  2. The LLM generates a structured Search Query (e.g., site:horizonorganic.com "Whole Milk" ingredients organic certification).
  3. A software wrapper executes this query via a Search API (like Google Search API or Bing Search API).
  4. The raw HTML or snippets from the top results are scraped and fed back into the system.

The difference LLMs may not actually be different; they might just be referring to different calls with different prompts and context. It’s probably better to have a smaller LLM generate the search query and a larger one process the results to generate a response.

“Products from different categories are characterized by different sets of uniquely defining attributes. For example, an alcohol product is uniquely defined by attributes such as vintage, aging, and flavor.”

If we’re storing these in a database, does it make sense to have a column for each possible field or use a non-relational database?

Not a good idea to have a column for every field.

You can also use a hybrid structure where you have a column in a relational database storing category-specific attributes in a JSON.

We can also have a tall table - Entity-Attribute-Value (EAV) Model - where a single product could have multiple rows.

“For each unannotated SKU, we first leverage OpenAI embeddings and the approximate nearest neighbors technique to retrieve the most similar SKUs from our golden annotation set.”

What exactly are they feeding into the embedding model? Also, is this just equivalent to dynamic few-shot prompting?

An SKU at DoorDash is represented by a list of attributes, including:

  • Item Name (e.g., "Dove Silk Glow Body Wash")
  • Size/UOM (e.g., "500 ml")
  • Brand (if available)
  • Category

To generate the embedding, DoorDash likely feeds a concatenated string of these available attributes (e.g., "Dove Silk Glow Body Wash 500 ml") into the OpenAI embedding model. They are embedding the raw merchant-provided text to find semantically similar items in their human-verified "golden" dataset.

And yes, this process is functionally equivalent to Dynamic Few-Shot Prompting.

How do you pick the right number of examples to provide as context for the LLM in the previous step?

Test out different thresholds (for number and for similarity), and evaluate metrics of interest: accuracy, latency/cost etc.

“Ultimately, the generated annotations are used to fine-tune an LLM for more scalable inference”

What does this mean?

Could mean training a distilled model on outputs of a larger teacher model. Scalability is improved both in terms of the new model being faster/cheaper but also no API limits if we use an open-source model.

How would I convince a business leader that this project was worth doing?

Any cost-savings are definitely worth highlighting.

For other impacts, they’d need to figure out a way to connect improvements in the performance of these intermediate steps to actual downstream user impact (ads CTR, search fulfilment rate, etc).

How would you improve this system if you were building it today?

  1. Use Contrastive Language-Image Pre-training (CLIP) to create unified embeddings that represent the text and the physical packaging in the same vector space.
  2. Graph-Augmented Retrieval (GraphRAG)
    1. Instead of just finding "items that look like this one" via vector similarity, GraphRAG allows the LLM to traverse the Knowledge Graph's relationships (e.g., "This item belongs to the same distributor as Brand X")
  3. Synthetic Data Generation
    1. Especially useful for edge-cases like misspellings.
2021 Blogpost original

Re-Scoring an ML Detection Engine on Past Attacks (part 1)

My notes

What I didn’t understand or am unsure about

Question

Answer

“Attackers are often even using ML systems of their own to adapt their attacks!”

How would attackers use ML systems? Would they run their planned attacks through ML models to see if they would be picked up and adjust important features if so?

That’s one way to do it. They can either build their own models or use any of the following:

  • Free/consumer filters — Gmail, Outlook.com — as a rough proxy.
  • Secure email gateways bought as a legitimate tenant (a few hundred dollars a month), giving unlimited test sends.
  • VirusTotal and public sandboxes, for the attachment/URL side.

There are other ways to use ML systems as well.

  1. Reconnaissance and targeting — classifiers over scraped LinkedIn/corporate data to rank who has payment authority, cluster orgs by vendor relationships, and time the attack to real events (an acquisition, a quarter close).
  2. Content generation — LLMs write fluent, on-register lures at scale, killing the grammar-error signal that older filters leaned on, and translate campaigns into new markets. Also style transfer: fed a few of a real executive's emails, mimic their cadence.
  3. Identity spoofing beyond text — voice cloning for the callback step of a BEC, deepfake video for "verify me on Zoom." This is the fastest-growing branch since 2023.
  4. Operational automation — ML to triage which compromised inboxes are worth working, and to run conversation threads with victims semi-autonomously.

“We cannot afford to throw out older attacks simply because they do not have the latest features. We must instead re-compute these features.”

Are there circumstances where it may be impossible to recompute those features?

Yes. Some examples include:

  1. URL detonation doesn’t work because page is no longer live
  2. Any internal features require underlying store to be append-only and timestamped
  3. TTLs for data (because of GDPR erasure requests, data-residency rules, PII minimization commitments, contractual deletion)

“This includes changes to underlying detection code, new datasets, new features, and the development of new models.”

What are examples of new data sets that would affect the pipeline? Everything else makes sense to me, but I think I need an illustrative example of datasets to understand it better.

The model: a message arrives with its own contents. A dataset is a side table keyed on something in that message — sender domain, sender address, recipient, URL, attachment hash — looked up and joined in before features are computed.

Four sources of such tables:

  1. The customer's own mail history. A per-tenant communication graph: for each sender-recipient pair, first-seen date and message count. Enables this vendor has emailed this AP clerk 400 times vs. never before. This class makes time travel hard — the correct join for a 2019 message is the graph as of 2019.
  2. The customer's other systems. A new Okta or HR-directory integration gives you the real reporting structure, so "CFO wires request to a direct report" separates from "CFO asks someone three orgs away."
  3. Third parties. A domain-intelligence feed keyed on sender domain, supplying registration date and hosting ASN — newly-registered-domain becomes computable the day you add it.
  4. Internal analyst or model output. A curated list of known-good vendor domains, or campaign cluster IDs from another job.

“we make an unintentional code change that modifies a feature used by a model (but we do not retrain the model)”

What is a likely example of such a code change?

Four classes:

  1. A refactor that alters a boundary case. Someone rewrites the sender-domain parser to handle subdomains properly. Correct by any reasonable spec — but a model that learned "this domain string is safe" now sees a different string on the same mail, and its lookup misses.
  2. A dependency upgrade. Bump the tokenizer, language-detection library, or HTML parser. Text features shift slightly across every message; no line of your own code changed, so nothing flags it.
  3. A change to a null or default. Missing sender-age gets -1 instead of 0, or a timeout that used to return "unknown" now returns "clean." The model learned the old sentinel's meaning and reads the new one as a real value.
  4. A units or normalization change. Timestamps switch from local to UTC, a count becomes a rate, a score gets rescaled to 0–1. The feature is now more sensible and entirely wrong relative to the trained weights.

Why rescoring catches these and unit tests don't: the bug is a mismatch between two artifacts — code and model weights — that no single component's tests can see. Running the whole stack over Golden Labels surfaces it as a drop in rescored recall, which is precisely the argument the post is making for testing the stack rather than the parts.

“Process to evaluate the entire detection system to produce performance analytics like precision, recall. We call this “rescored recall.””

This isn’t equivalent to actual recall, right? Because they don't have all true positives.

Yes, it’s not equivalent.

“We must have correct and unbiased features”

The word "unbiased" implies the existence of biased features. Is this just a reference to time travel?

Yes, I think so.

“Should be able to update all historical samples”

Let's say the data you have doesn't let you update all historical samples, but only a subset. Do you just not use it then? That seems a little inefficient.

Not sure how Abnormal handles this internally but some alternatives to throwing it out include:

  1. Mark the feature missing and let the model handle it. Tree models natively route missing values; for others you carry an explicit is_null indicator. Cost: the model learns a pattern from missingness itself, and if missingness correlates with age, it correlates with label. Mitigation is to inject the same missingness into recent samples at the observed rate so the indicator carries no time signal.
  2. Restrict the metric's window, keep the sample. Report rescored recall on the full corpus for stack-wide features, and on the post-cutoff window for the new feature's contribution. The old attacks still train models and still count in every other evaluation. This is usually the right answer for a feature added six months ago.
  3. Two-tier the corpus. A fully-hydrated recent tier for anything requiring complete features, plus a degraded historical tier used where partial features suffice. Explicit, but adds bookkeeping to every consumer.
  4. Impute. Almost always the worst choice here. An imputed value on a rare attack is a fabricated signal on your most precious samples, and it's invisible downstream — unlike a null, which announces itself.

“Time travel ensures unbiased data and also helps us avoid future leakage: avoiding bringing any data from the future (to a past sample) that leaks labels into training features.”

What are a couple of examples of features that could be affected in this way?

  1. Sender-recipient relationship count. A behavioral feature like "number of prior messages from this sender to this org." Recomputed today over the full history, an attacker domain that ran a 2019 campaign shows hundreds of messages — inflating it into a familiar sender. Worse, the count includes the campaign's own later messages, which exist because the attack happened. Time-traveled to the attack timestamp, the correct value is zero or near-zero, which is the actual signal.
  2. Domain reputation from a threat feed. Query the feed today for a domain used in a 2019 attack and it returns "known malicious" — because the industry blacklisted it in response to that campaign. That feature is a near-perfect label proxy in training and worthless at inference, where the domain is hours old and unlisted. A model that leans on it will look excellent on Golden Labels and fail live.

“To do so, we rely on the most recent updated Golden Labels from the night before”

Is there a reason they're doing daily batching? Couldn't this miss recent attack patterns?

Why daily rather than continuous:

  1. Cost. Full regeneration re-extracts features for every historical sample against time-correct join tables — a Spark job over years of data. Running it hourly multiplies compute by 24 for a corpus that changes by a fraction of a percent.
  2. Labels aren't available in real time anyway. A message detected today may not be labeled for hours or days — customer reports, analyst review, retroactive discovery. A faster pipeline would mostly re-process unlabeled data.
  3. Reproducibility. A dated, branch-tagged snapshot means two engineers comparing baseline vs. experiment are measuring against identical data. A continuously-updating corpus makes results non-comparable across runs.
  4. Off-peak windows. Overnight batch avoids contending with production scoring for cluster capacity.

On your concern — the gap is real but it's not where the risk lives. Two things to separate:

  • Detection latency — how fast a new attack pattern gets blocked. Not affected by this pipeline. That runs through the live scoring path and, presumably, faster mechanisms: signature/campaign clustering, analyst-pushed rules, threat-intel updates. Rescoring is an offline evaluation and training-data system.
  • Evaluation and retraining latency — how fast a new pattern shows up in your metrics and your next model. Bounded at ~24 hours by this design, plus label delay, which likely dominates.

“We can either set up two configurations, “Baseline” and “Experiment,” or run this on two different code branches.”

What are the trade-offs between these two options?

What each can express

  • Two configs, one branch. Only changes the config schema anticipated: swap a model, skip a stage, point at a different artifact. Anything requiring new code is out of reach.
  • Two branches. Anything at all — new feature logic, changed extraction code, a new join, a restructured decision tree. This is the only option when the change is code, which is exactly the failure mode the post worries about in the "unintentional code change" paragraph.

Trade-offs

  1. Confounding. Configs vary one declared knob against identical code — a clean A/B. Branches carry every difference between them, including drift from main if the branch is stale. Diagnosing "which of my six commits moved recall" is on you.
  2. Compute. Configs can share upstream stages: run feature extraction once, fork only at MODEL_SCORING. Two branches generally can't share intermediates, since you can't prove the shared prefix is identical — so you pay roughly double.
  3. Reproducibility. A config is a serializable object you can attach to a result and re-run later. A branch is a moving pointer; rebasing or force-pushing silently invalidates the comparison unless you record the commit SHA.
  4. Path to production. A config swap is an experiment that then needs code written to ship it. A branch experiment is already the artifact you'd merge, so the thing you measured is the thing you deploy.

Rule of thumb: configs when the change fits the schema — cheaper, cleaner, more reproducible. Branches when it doesn't, accepting more compute and more attribution work. The post's stated aspiration, CI/CD auto-running rescoring on affected stages per code change, is the branch path automated.

2024 Blogpost original

Building Confidence: A Case Study in How to Create Confidence Scores for GenAI Applications

It’s straightforward to ask GenAI tools to extract information from documents but confidence scores for these predictions are necessary to support human-in-the-loop decisions and meet regulatory requirements.

My notes

Summary

Context

The Financial Engineering team at Spotify aims to enhance efficiency in financial processes by automation. One of the tasks they work on is extracting information from invoices using GenAI.

Problem

It’s straightforward to ask GenAI tools to extract information from documents but confidence scores for these predictions are necessary to support human-in-the-loop decisions and meet regulatory requirements.

Solution

They considered three approaches:

  1. Separate calibrator model
  2. Averaged Logprobs of outputs
  3. Majority voting amongst different models

The last approach was found to be most effective

Some key insights/new things I learned

  1. Majority voting amongst different models can be an effective way to produce confidence scores for GenAI applications
  2. Prompt permutation can increase effective sample size for confidence score generation

Design Details

  • Data
    • Invoices (details not specified)
  • Architecture
    • Separate calibrator model
    • Averaged Logprobs of outputs
    • Majority voting amongst different models
      • Weighted models based on accuracy
      • Platt Scaling for calibration of confidence scores
  • Evaluation
    • Metric(s)
      • Plotted graphs and computer correlation of accuracy vs confidence scores
    • Data
      • Invoices (details not specified)
    • Results
      • Majority voting performed best, further improving after calibration

What I didn’t understand or am unsure about

Question

Answer

“In many cases, GenAI-powered solutions are easier to maintain than traditional ML models”

Why is that?

  1. The artifact you maintain is a prompt, not a training pipeline. A traditional supervised model for invoice parsing needs labeled data, a feature pipeline, a training job, a model registry, and a serving stack. Each is a thing that breaks. With GenAI the vendor owns the model; you own a text file and some post-processing.
  2. Change cost is asymmetric. A new vendor with an odd invoice layout means, for a traditional model, collecting and labeling examples of that layout, retraining, revalidating, redeploying — days to weeks. For GenAI it means editing the prompt and rerunning evals — hours. This is the maintainability consequence of the flexibility bullet.
  3. No feature engineering to keep alive. Document parsing models traditionally depend on OCR output, bounding-box geometry, layout heuristics, regexes per field. Those are brittle and accumulate special cases. An LLM absorbs that into the model's pretrained understanding.
  4. No label drift treadmill. Traditional models degrade as the input distribution shifts and need periodic relabeling and retraining. GenAI doesn't require you to own that loop.
  5. Fewer specialists needed. A prompt plus evals is legible to an engineer or analyst; a fine-tuned layout model needs someone who understands its training setup.

Why didn’t they ask the models making the predictions to generate confidence scores? Even if those scores are likely to be over-confident, that feels like something that can be addressed through calibration.

There are a few reasons why this might not work great (lack of independence, instability across runs, interpretability) but I think it’s still worth at least discussing.

““Calibrator” refers to using a separate single GenAI model to evaluate the outputs generated by other models and assign confidence scores”

  1. Should this be a different model than those used for extraction? If so, which ones make sense for which use-case?
  2. What would actually be included in the prompt for this model?

1. Should it be a different model?

Yes, for the same reason self-scoring fails: correlated errors. If the extractor and the judge are the same model, a misread "0" as "O" looks correct to both. A different model gives you an independent error distribution — the same property that makes their voting ensemble work.

Which model:

  • Judging structured extraction against a document (their case) — needs strong document/vision grounding, since the judge must re-read the invoice image, not just inspect the text. That points to a frontier multimodal model, and specifically not the weakest member of the extraction ensemble.
  • Judging free-text output with no ground-truth artifact — a cheaper text-only model is usually fine, since the task is coherence and consistency rather than verification.
  • High-volume pre-filtering — a small fast model as a first pass, escalating only the uncertain cases.

2. What goes in the prompt?

Six things:

  • The source document (image or OCR text) — without it the judge is grading plausibility, not correctness.
  • The field definition, ideally the same one the extractor got, so both are answering the same question about "invoice total."
  • The candidate value to be judged.
  • A rubric that decomposes the judgment: is the value present in the document, is it in the right place, is the format right, is there a competing candidate elsewhere on the page.
  • A required structured output — score plus a short justification, ideally with the justification first so the score follows from stated evidence rather than being emitted cold.
  • A discretized scale with anchored labels (e.g. five bands, each defined) rather than a free 0–100. LLMs cluster at round numbers, and defined bands are more stable across runs — which was Spotify's stated complaint about calibrators.

“The calibrator model is also trainable by learning from feedback and can potentially improve over time.”

Does this mean fine-tuning, few-shot learning, prompt iteration or something else?

Not stated in the blogpost but my best guess would be few-shot examples plus prompt iteration, not fine-tuning.

The reasoning: their whole framing is that GenAI wins on fast development and low maintenance. Fine-tuning a judge would undercut both — it needs a labeled dataset of (document, extraction, correct-or-not) pairs, which is the labeling burden they were trying to escape by not training an extractor in the first place. And the invoice distribution keeps shifting as vendors change, so you'd be retraining continually.

“confidence scores generated by calibrators are difficult to interpret and sometimes counterintuitive”

In what ways is this true?

Difficult to interpret — the number has no defined referent. A voting score of 80% means four of five models agreed; you can state what generated it and reproduce it. A calibrator's 0.8 means whatever the model's next-token distribution landed on. It isn't a frequency, isn't tied to an observable event count, and can't be explained to an auditor as a control.

Counterintuitive — the specific ways this shows up:

  • Non-monotonicity. Two extractions where one is clearly harder get similar or inverted scores. The judge is responding to surface features (clean formatting, plausible-looking numbers) rather than to actual difficulty.
  • Clustering. Almost everything lands at 0.85–0.95, so the score doesn't separate the cases you care about. Then an occasional 0.3 appears for no visible reason.
  • Score-justification mismatch. The judge writes a reason that identifies a real problem, then attaches a high score anyway — the text and the number are generated somewhat independently.

“the scores are not consistent across multiple runs”

What’s the underlying reason for this and are there any good ways to address it?

Underlying reason. Sampling is the obvious part — temperature above zero means the score token is drawn from a distribution, and if that distribution has mass spread across "8", "85", "9", you get different numbers on identical inputs. But the deeper issue is that the underlying distribution is flat. A well-grounded judgment concentrates probability on one answer and survives sampling; a weakly-grounded one doesn't. The instability is a symptom that the judge has no firm basis for the number, not just that decoding is stochastic.

Two secondary causes: LLMs are sensitive to irrelevant perturbations (ordering, whitespace, which examples are in the few-shot block), and even at temperature zero, batching and hardware nondeterminism in floating-point reduction make outputs non-reproducible run to run.

What actually helps:

  • Discretize the output. Five anchored bands instead of 0–100. Sampling noise that would move 87 to 91 usually doesn't cross a band boundary. Biggest gain for the least effort.
  • Sample it repeatedly and aggregate. Run the judge n times and take the median or modal band. This converts an unstable point estimate into a stable one — and note it's the same move as majority voting, which is where they ended up.
  • Force the reasoning before the score, so the number is conditioned on stated evidence rather than emitted cold. Reduces variance but doesn't eliminate it.

“However, the methodology is not always transparent – it varies by provider, and for many models it is unknown how they were calculated.”

Don’t they all work the same way? If not, what are the differences?

The math is identical everywhere — log of the softmax over the vocabulary at one decoding step. What varies is everything between that number and what reaches you.

The three that actually matter for Spotify's method:

  1. Availability. Not a uniform feature. Anthropic doesn't support logprobs at all; Gemini's support lives on Vertex AI rather than the standard API. A 2025 study of OpenRouter found only 164 of 813 endpoints (23%) returned them. A multi-vendor ensemble can't get logprobs from every member — which alone may explain why they dropped the approach.
  2. Transformations applied before the number you see. Temperature scaling, top-k/top-p truncation with renormalization, repetition penalties, logit bias, and constrained JSON decoding all alter the logits first. Constrained decoding is the worst offender here: zeroing invalid tokens inflates confidence in whatever survives. Providers rarely document which of these apply.
  3. Tokenization. Averaging over tokens makes the denominator tokenizer-dependent. "1,234.56" might be 3 tokens in one model and 8 in another, so the same extracted value gets a different average across ensemble members and isn't comparable.

Points 2 and 3 are what break their specific method — exponentiating an average of logprobs assumes a clean posterior, and under those transformations it isn't one.

“The confidence score was then calculated by exponentiating the average logprobs of the tokens.”

Why exponentiated?

To get back to a probability scale.

Figure from the post

What does a single dot here represent?

Almost certainly a bucket of extractions grouped by confidence score, not an individual extraction. Two reasons:

  • The y-values are accuracies — 60%, 75%, 96%, 22%, 12.5%. A single extraction is right or wrong, so accuracy is only definable over a group.
  • The specific fractions give away small bucket sizes: 60% = 3/5, 75% = 3/4, 12.5% = 1/8, 66% = 2/3, 25% = 1/4. Roughly 8–12 dots per panel, each holding a handful of invoices.

“Majority voting is an ensemble method that selects the final output by choosing the most common response from multiple prompts or among multiple GenAI models.”

Are they using majority voting for the actual prediction too or just for the confidence scores? If it’s only for the confidence scores, couldn’t you potentially have the single extractor model have a different output than the majority vote?

I think they’re using it for both, which means the concern I raised is not applicable.

What models are generally considered best for this kind of use-case?

Frontier multimodal: Gemini 3 Pro led March 2026 invoice benchmarks (94.75%), with Azure Document Intelligence (90.52%) and Claude Sonnet 4.5 (90.27%) behind it. Claude Sonnet 3.5 degraded least as document quality dropped from clean PDFs to photos — resilience matters more than peak accuracy for real vendor mail.

Open-weight: Qwen3-VL-8B-Instruct is now state-of-the-art among open vision-language models.

Specialized engines (Azure, ABBYY, Rossum) remain competitive, particularly on tables and line items.

Two things that matter more than the rankings:

  1. Pipeline structure beats model choice. A 2026 study on scanned financial documents found feeding full PDFs directly to a VLM was the worst configuration in every test; OCR-first multistage pipelines won by a wide margin. The document-specialized MiniCPM-o-2.6 also beat the much larger general-purpose Gemma-3-27b under identical settings.
  2. For the ensemble, diversity beats individual accuracy. The voting signal comes from independent error distributions, so mix vendors and architectures. Logprob support is irrelevant here, so Anthropic's lack of it doesn't rule Claude out the way it would for the logprob approach.

“Literature often suggests using four to seven models in an optimal ensemble.”

What literature?

I haven’t found anything that supports this specific number so far so more extensive research may be needed.

In their use-case, they found that majority voting works best. Is that consistent with what the literature on this has found more generally?

Where the literature agrees. Consistency-based confidence is the mainstream approach, resting on Spotify's same premise: genuine knowledge produces reproducible answers, fabrications diverge. A recent head-to-head audit found self-consistency had the highest rank correlation with correctness on GPQA (ρ=0.31), ahead of verbalized confidence (0.21) and P(True) (0.19) — matching their ordering.

Their logprob failure also replicates, with a better diagnosis than the post offers. A 2026 paper on document field extraction notes these methods measure confidence in the generated token sequence, but extraction errors come from what the model can't observe — unreadable source material, ambiguous layouts, OCR noise. A model confidently transcribing OCR noise emits high logprobs for a wrong answer.

Where it's less flattering. That paper puts self-consistency across five calls at only 0.744 AUC on field extraction — better than logprobs, but not good, and at five times the cost. The audit concurs that no signal was a reliable standalone abstention score.

One way Spotify did better. Most of this literature tests self-consistency, resampling one model. They used cross-model voting, drawing on genuinely independent error distributions rather than one model's sampling noise — strictly stronger, and worth weighing against that 0.744.

The frontier is multi-signal. The 2026 paper argues for combining agreement with document-side evidence: image quality, spatial location, OCR text. That directly addresses the long-text-field problem they left unsolved.

“To calculate the final score, a weighted majority voting approach was implemented. Weight for each model was based on model accuracy and then normalized to sum to one.”

Do we need any kind of train-test split here?

Yes, and in two places.

1. The weights. "Model accuracy" has to be measured on labeled invoices. If you measure it on the same invoices you then report ensemble performance on, the weights are tuned to that set's noise — models that happened to get lucky on the hard cases get upweighted, and reported accuracy is optimistically biased. With only 5–6 weights the overfitting is mild, but it isn't zero, especially if the eval set is small (their Figure 1 buckets suggest it is).

2. Platt scaling — this is the bigger one. Platt scaling fits two parameters mapping raw confidence to calibrated probability, and it must be fit per field. Evaluate calibration on the fitting data and your ECE looks great by construction. Their Figure 3 shows before/after calibration curves; if those are in-sample, the improvement is partly guaranteed rather than measured.

What the setup should be: three splits, or nested cross-validation. One fold to estimate model accuracies for weights, a second to fit Platt parameters, a third untouched fold to report accuracy and calibration.

Why it matters more here than usual. The confidence score is a control — a threshold decides whether an invoice goes to a human. If calibration is optimistic in-sample, the real pass rate at a 95% threshold is worse than the number they'd show an auditor. Under SOX, an overstated control is the failure mode that actually costs you.

“we applied Platt scaling”

Why Platt scaling as opposed to isotonic regression?

Three reasons, in order of weight:

  1. Sample size. Platt fits two parameters; isotonic fits a free monotone step function and needs roughly a thousand-plus labeled examples per field to avoid overfitting. They calibrate per field, and their plots suggest far less data than that.
  2. Input granularity. With 5–6 models, majority voting yields only about six distinct confidence values. Isotonic would memorize an accuracy per value — effectively a lookup table with no ability to interpolate. A sigmoid smooths across levels.
  3. Auditability. They chose linear weights over exponential explicitly for explainability. Two parameters is easier to document to an auditor than a nonparametric fit, and a sigmoid's smooth output avoids the step discontinuities that would make a threshold-based control jumpy.

Where isotonic would win: miscalibration that isn't sigmoid-shaped, such as good calibration mid-range but distortion only at the top end. Platt can't represent that — but at their sample sizes they likely couldn't detect it either.

“This approach requires careful tuning of distance thresholds.”

What would this tuning look like concretely?

Concretely, three separable pieces:

1. What you're tuning. A cosine-distance cutoff τ: two model responses join the same cluster if distance < τ. Too low and genuine paraphrases ("123 Main St." vs "123 Main Street") split into separate clusters, so agreement is understated and confidence collapses. Too high and real disagreements merge — the "0" vs "O" failure they hit — so confidence is overstated on a wrong answer. Related knobs, if you're being thorough: the embedding model, and the linkage rule (single-link chains transitively and will bridge distinct answers).

2. The objective. Not clustering quality. Build a labeled set of response groups where you know which responses are equivalent, then sweep τ and score against the two error types above. Because false merges are the dangerous direction in finance, weight precision over recall — e.g. maximize recall subject to false-merge rate under some ceiling — rather than optimizing F1.

3. Per-field fitting. One τ won't serve addresses, item descriptions, and vendor names — their embedding-distance distributions differ. That means a labeled set and a sweep per field, plus a held-out fold, since τ is a fitted parameter like the Platt coefficients.

Why it still failed for them: embedding distance is dominated by semantics, and "0" vs "O" is semantically invisible. No τ separates that case, which is why this is a threshold problem you can't fully solve by tuning the threshold.

“This approach requires extensive prompt engineering to ensure the selector prioritizes factual accuracy and consistency.”

How does factual accuracy fit into this? From my understanding, the selector model is just looking at the responses of the voter models, right? What facts are being verified?

Unclear from the post.

“we broke down long text fields into smaller, more manageable components (e.g., splitting addresses into street, city, state, and zip code)”

How did they do this break-down? Using a model or heuristic/regex?

Two possible readings, and the wording favors the first:

  1. Split at the prompt, not after. Each model is asked for street, city, state, and zip as separate fields, so five separate votes happen on five short strings. Nothing needs parsing — the decomposition lives in the extraction schema. This fits "broke down long text fields into smaller components," fits their stated goal of improving agreement during voting, and fits their whole approach of solving problems by prompt adjustment rather than added machinery.
  2. Split after extraction, parsing each model's full address string into components before voting. This is strictly worse for their purpose: if the parser is a regex it breaks on international addresses, which is their stated problem domain; if it's a model, you've added a component whose own errors now contaminate the vote.

Should we be concerned about the intra-model correlation in the prompt permutation phase? A 94% score suggests a high degree of confidence but if the same models are being used, I feel like it’s somewhat exaggerated.

I think this is a reasonable concern but maybe mitigated somewhat by the calibration stage.

How would I convince a business leader that this project was worth doing?

Hard to say without more insight into the use-cases of human-in-the-loop decisions and meeting regulatory requirements.

How would you improve this system if you were building it today?

1. Add document-side signals to the score

The biggest weakness is that every signal they use is model-side. Agreement can't detect a failure all models share, and their unsolved "0" vs "O" case is exactly that. Current work on extraction confidence combines vote agreement with evidence from the document itself: image quality metrics, whether the extracted value is spatially locatable on the page, agreement with an independent OCR pass, and agreement between two structurally different extraction calls. A wrong character is invisible to voting but visible to a string comparison against OCR text.

2. Use business validation as a free, hard signal

Invoices are internally redundant in ways they never exploit: line items should sum to subtotal, subtotal plus tax should equal total, currency should match the header, the invoice number should match the PO's, the vendor should exist in the master file. These are deterministic checks with no false-confidence failure mode. A total that fails arithmetic should be low-confidence regardless of unanimous agreement, and a total that passes deserves a lift. This is cheaper than any additional model run.

3. Restructure the ensemble spend

Instead of running 5–6 frontier models on every invoice: run 2–3 on the first pass, and escalate to the full ensemble plus permutation only when the cheap signals disagree or validation fails. Most invoices are native digital PDFs where accuracy is 98–99%; the cost belongs on the photographed and multilingual tail.

Two related fixes: two-stage voting for the permutation phase (per-model modal answer, then vote across models) so correlated prompts don't inflate the denominator; and mixing architectures deliberately — a specialized engine like Azure alongside frontier VLMs — since diversity is what the vote is actually measuring.

4. Fix the pipeline, not just the aggregation

A 2026 study on scanned financial documents found feeding full PDFs directly to a VLM was the worst configuration in every test; OCR-first multistage pipelines won by a wide margin. If they're still passing whole documents, upstream structure will buy more than another voter.

5. Tighten the evaluation

The post shows no train/test discipline for the weights or the Platt parameters, and the calibration curves may be in-sample. Nested splits, reported AUC and AURC rather than just correlation, and a stated false-merge rate — because under SOX an overstated control is the failure that costs you.

2022 Blogpost original

Evolving DoorDash’s Substitution Recommendations Algorithm

When a product in an order is out-of-stock, DoorDash can either skip that item or find a substitute. Skipping is a bad user experience. Finding a substitute manually can be a lot of work (the Dasher has to contact the user).

My notes

Summary

Context

Sometimes the products users order on DoorDash are out-of-stock.

Problem

When a product in an order is out-of-stock, DoorDash can either skip that item or find a substitute. Skipping is a bad user experience. Finding a substitute manually can be a lot of work (the Dasher has to contact the user).

Solution

Substitution recommendations algorithm built in three stages:

  1. Unsupervised
  2. Boosted trees
  3. Deep learning

Some key insights/new things I learned

  1. Launching an unsupervised model can be a good way to handle cold-start and collect data for supervised models down the line

Design Details

  • Data
    • item metadata
    • “catalog team built out a well-defined taxonomy”
    • After launching the initial unsupervised model, they asked consumers to rate suggested substitutions as thumbs up or thumbs down.
  • Architecture
    • Stage 1 (unsupervised)
      • TF-IDF cosine similarity based on an item’s name
      • Well-defined taxonomy that let us apply heuristics on top of the text-based similarity score to restrict recommendations to relevant categories
    • Stage 2 (supervised, boosted trees)
      • LightGBM binary classifier
    • Stage 3:
      • Deep learning recommendation model
      • “Categorical features (or in this context, items in our catalog) are processed as embeddings and there is a bottom MLP that encodes our dense feature.” “Item embeddings trained on the search behaviors of DoorDash users”
      • “Next, feature interactions are computed explicitly and the results are processed to discern a top MLP, which is fed into a Sigmoid function to yield a probability score”
  • Evaluation
    • Data:
      • “While we were using an unsupervised model, we leveraged manual reviews to measure recommendation quality. That involved identifying top-selling items across product categories and curating ideal substitutions for them to create a “golden” dataset.”
    • Metric(s):
      • Unsupervised: compared what percentage of the algorithm’s recommendations were matched by human-curated substitutions
      • Supervised: AUC, customer approval rate, coverage (percent of ordered items with recommendations)
    • Results: not specified
  • Deployment
    • Not specified
  • Post-deployment
    • Not specified

What I didn’t understand or am unsure about

Question

Answer

Does the model run on all items the user adds to their cart or just the ones that turn out to be unavailable?

For all ordered items, not just the ones that turn out to be unavailable. Serving is almost certainly a precomputed item→candidate-substitutes table refreshed in batch, not a live forward pass over the whole catalog per cart.

“then proceeded to binary classification, and eventually pursued a deep learning recommendation model”

The way this is phrased makes it seem like the deep learning model isn't doing binary classification but something else instead?

The phrasing is imprecise - the DLRM is still doing binary classification.

Why did they use TF-IDF cosine similarity as opposed to semantic embeddings (like BERT) at the start?

No reasons presented in the blogpost but possible reasons include:

  1. Off-the-shelf BERT is a poor fit for product titles specifically. Grocery SKU names are short, keyword-dense strings with brand tokens, sizes, and units — "Coca-Cola Classic 12 fl oz Can 12pk." Generic BERT is pretrained on prose and, without fine-tuning, its sentence embeddings underperform even simple lexical baselines on this kind of text; it also tends to wash out the exact-token matches (brand, pack size) that carry most of the signal here. TF-IDF preserves those.
  2. Operational cost. TF-IDF is a sparse matrix and a dot product — cheap to compute over a large catalog, trivial to refresh when SKUs change, and fully inspectable when a recommendation looks wrong. A transformer at phase 1 means GPU inference, an embedding store, and a debugging surface you can't eyeball. For an MVP whose main job was to unblock feedback collection, that overhead buys little.

Why did they only use the item’s name as opposed to descriptions to find similar items?

No reasons presented in the blogpost but possible reasons include:

  1. Descriptions largely didn't exist. DoorDash's grocery catalog was assembled from merchant feeds and in-store data collection, not authored by DoorDash. Names are the one field every SKU has; descriptions are missing, truncated, or inconsistent across merchants for a large share of items. A feature you can only compute for part of the catalog directly caps the coverage metric they later tracked.
  2. For groceries, the name carries nearly all the discriminative signal. "Bush's Best Cut Green Beans 14.5 oz Can" already encodes brand, product type, form, and size — the exact axes that determine a good substitute. This is unusual: for restaurants or apparel, descriptions matter much more. Grocery SKU names are effectively structured records rendered as strings.
  3. Descriptions would have actively hurt a TF-IDF model. They're mostly marketing boilerplate and shared brand copy. Under bag-of-words cosine similarity, two unrelated products from the same manufacturer share long stretches of identical text and score as similar. IDF down-weights terms common across the whole catalog, but not terms common within a brand or category — so the noise survives. Longer documents also dilute the weight of the few tokens that matter (the pack size, the brand).

“our catalog team built out a well-defined taxonomy that let us apply heuristics on top of the text-based similarity score to restrict recommendations to relevant categories”

I didn't fully understand this. What is an illustrative example?

What it means: two filters in sequence. TF-IDF says "these names share words." The taxonomy says "these are the same kind of thing." Rank by similarity within a category node; drop anything outside it regardless of score.

Example. Customer orders Coca-Cola 12-pack cans. TF-IDF alone surfaces Pepsi 12-pack and Diet Coke 12-pack (good), but also Coca-Cola Gummy Candy, Coca-Cola BBQ Sauce, and a Coca-Cola T-Shirt — all scoring high on the brand token. The taxonomy cuts the last three: wrong category node.

“chose to use LightGBM for this phase because of both its relatively high performance with minimal hyperparameter tuning”

High performance as measured on their training data or based on prior literature?

Unclear from the blogpost. It’s worth noting that "performance" here could mean predictive accuracy or computational efficiency. LightGBM's actual differentiator versus XGBoost is speed — histogram-based splitting and leaf-wise tree growth — with accuracy roughly comparable

“this model combines principles from approaches based on collaborative filtering and predictive analytics”

In what ways does this model combine those principles?

Collaborative filtering side = the embeddings plus the explicit interaction step. Learning a vector per item and scoring pairs by dot product is matrix factorization, the canonical CF method. DLRM computes pairwise dot products among all embeddings, so that layer is essentially a bank of MF models. The signal is behavioral: which items co-occur or get approved together, with no notion of what the item is.

Predictive analytics side = the two MLPs over dense features. Continuous features (price, size, popularity, similarity scores) go through the bottom MLP and, after interaction, the top MLP — a standard supervised classifier on engineered features, which is exactly what phase 2's LightGBM was. The signal is attribute-based.

Why combining matters here: pure CF fails on cold-start items with little feedback; pure feature-based models miss preference patterns not encoded in attributes (the Pepsi 12-pack over 2-liter Coke case). DLRM lets both flow into the same sigmoid.

One caveat: DoorDash's item embeddings were pretrained on search behavior, not learned fresh from substitution labels. So the CF component partly comes in as a frozen input rather than being fit end-to-end — the post doesn't say whether those embeddings were fine-tuned.

“categorical features (or in this context, items in our catalog) are processed as embeddings”

How are items represented in this model? A single embedding per item or broken down into different features (some of which would be categorical features represented as embeddings)?

Most likely a single vector per item, not a decomposition into brand/size/category sub-embeddings. The post's own phrasing points this way, and the semantic embeddings it inherited are whole-item vectors by construction. Attribute-level features are named as future work ("organic," "kosher," image embeddings), which implies they weren't decomposed at the time. Standard DLRM would allow additional categorical fields — taxonomy node, brand, merchant/store, each with its own table. Plausible here given the catalog taxonomy already existed from phase 1 but bot mentioned either way.

Note that we have at least two item embeddings per scoring pass. The task is pairwise — score candidate j as a substitute for ordered item i. So both IDs get looked up, and the explicit interaction layer's dot product between them is where most of the signal lives. The post never says this, but the task formulation requires it.

“Next, feature interactions are computed explicitly”

What does this mean?

In DLRM, "explicit interaction" is a specific operation: take every embedding vector plus the bottom MLP's output, and compute the dot product between every pair. Those scalars get concatenated with the dense vector and passed to the top MLP.

"Explicit" contrasts with "implicit." If you just concatenate all inputs and feed an MLP, the network can in principle learn multiplicative relationships between features, but it has to discover them through stacked additive layers — slow and sample-hungry. DLRM instead hard-codes second-order terms, borrowing from factorization machines. Here that's very natural: the dot product between the ordered item's embedding and the candidate's embedding is a similarity score, which is the core signal for substitution.

What is the difference between the bottom MLP and the top MLP?

Figure from the post

Why the bottom MLP exists at all (inferred): dimensional compatibility. The interaction layer takes dot products between vectors, so every participant must have the same width. Item embeddings might be 64-d; the raw dense features are some arbitrary count of scalars (price, size, popularity, TF-IDF similarity). The bottom MLP projects them to 64-d so the dense side can participate in the same pairwise dot products as the items. Without it, dense features couldn't interact with embeddings at all — you'd be stuck concatenating them at the end.

Why the top MLP exists: the interaction layer outputs a bag of scalars with no ordering or weighting. Something has to learn which of those pairwise agreements matter and how they combine nonlinearly. That's the top MLP — it's the actual classifier, and it's where higher-order effects beyond second order get modeled.

Rough intuition: bottom MLP is a translator getting dense features into embedding space. Top MLP is the judge deciding the verdict from all the evidence.

“the embeddings are trained on the search behaviors of DoorDash users”

In 2-3 sentences, explain how the search behavior is used to train the embeddings.

Search logs supply pairs of things users treated as equivalent — a query and the item they clicked, or two items clicked under the same query. A twin (Siamese) network encodes each side with shared weights and is trained with a contrastive objective: pull matched pairs together in vector space, push randomly sampled non-matches apart. The result is that items people search for interchangeably end up nearby, which is why green beans land closer to peas than to corn.

“For example, as shown in Figure 4, the LightGBM model recommended canned corn as a substitute for canned green beans. The deep learning model, however, recommended canned green peas, because item embeddings accurately represent that beans are more similar to peas than corn.”

Is there something intrinsic to these models that leads to LightGBM performing worse for this example or is it more random?

It's intrinsic — but the cause is representation, not "trees are worse than neural nets." Two distinct mechanisms:

1. Nothing in LightGBM's feature set separates corn from peas.
Trees split on features you hand them. Under plausible phase-2 features — name token overlap, taxonomy node, price, size, popularity, per-pair feedback rate — canned corn and canned green peas are near-identical for a green-beans query. Same node (canned vegetables), same price band, same "canned" token. With no discriminating feature, the model falls back on the strongest remaining signal: popularity. Canned corn outsells canned peas, so it wins. That's a systematic bias, not noise.

2. Even given embeddings, a tree can't cheaply express pairwise similarity.
Feed LightGBM the 64-d item vectors as columns and it still can't compute a dot product — axis-aligned splits approximate a similarity surface only with an enormous number of splits. DLRM's explicit interaction layer does it in one operation. This is a genuine architectural gap for pairwise tasks.

“We then compared what percentage of the algorithm’s recommendations were matched by human-curated substitutions.”

Is this equivalent to precision, recall, or something else?

Precision@k

“coverage (percent of ordered items with recommendations)”

Shouldn’t this be 100% by construction?

No — and the reason is that the model was deliberately allowed to abstain.

A pairwise scorer over the whole catalog will always return something if you take the argmax. But shipping the argmax regardless of score is bad: a wrong substitution is worse for the customer than no suggestion, since the fallback (chat with the Dasher, or a refund) already exists and is acceptable. So there's a confidence threshold, and items whose best candidate falls below it get nothing.

“We also plan to develop personalized recommendations because we’ve observed that customers have highly individualized substitution preferences”

What would be the best way to incorporate personalization into this system?

Three places personalization can enter, by pipeline stage:

1. Input layer — a user ID embedding as an additional categorical field.
DLRM-native: add a user embedding table, and its dot products with the ordered-item and candidate-item embeddings fall out of the existing interaction layer for free. Architecturally the cleanest. Directly hits the sparsity wall above, so it only works for the head of the user distribution.

2. Feature layer — user-level dense features into the bottom MLP.
Engineered signals that generalize across users rather than identifying them: this customer's historical brand-switch rate, private-label acceptance rate, size-flexibility rate, organic-purchase share, price sensitivity, past thumbs-down rate by category. Each is estimable from a handful of events and degrades gracefully to a population prior when data is thin.

3. Post-scoring layer — constraints and re-ranking on a global model.
Hard filters from stated preferences (dietary restrictions, allergen avoidance, no-substitute flags) plus boosts from the customer's own purchase history — a candidate they've bought before should outrank an equally-scored one they haven't.

How would I convince a business leader that this project was worth doing?

  1. Come up with some estimate of the cost to DoorDash of
    1. Users not receiving the item they wanted because out of stock
    2. Friction from manual substitution
  2. Calculate how much of that cost has been mitigated through the algorithm

If you’re A/B testing, you can also look at various user engagement and revenue metrics

How would you improve this system if you were building it today?

Ranked by expected value.

1. Condition on real-time store-level availability.
The system as described scores substitutes for an item in the abstract. But a recommendation is worthless if the suggested item is also out of stock in that store, and the Dasher is standing in the aisle with a list. I'd make the candidate set store-and-time-specific: score P(good substitute) × P(in stock at this store now), where the second term comes from an availability model fed by Dasher found/not-found reports, POS feeds where available, and recency of last confirmed pick. This is the largest gap in the original design and it's orthogonal to model quality — a mediocre recommender that only suggests things actually on the shelf beats an excellent one that doesn't.

2. Generate the missing metadata with a VLM instead of waiting for merchants.
The post's own next steps name richer attributes (organic, kosher) and image embeddings as future work, and my earlier answer on why they used names only came down to descriptions being absent. That constraint is largely gone. Run a vision-language model over product images and packaging text to extract a structured attribute record per SKU: brand, form (canned/frozen/fresh/dried), net quantity and unit, flavor, dietary flags, allergens, private-label indicator. Those become explicit features rather than something the model has to infer from a name string, and they make the failure modes debuggable — you can point at why a swap was blocked.

3. Fix the label.
Thumbs-up/down on a shown recommendation is a preference proxy collected before the fact, and it's only observed for pairs the system chose to display — classic selection bias, where the model is trained on its own past decisions. I'd train on realized outcomes instead: was the substitute accepted, kept versus refunded or returned, and how was the order rated. Combine with logging propensities so offline evaluation can use inverse-propensity or doubly-robust estimators rather than accuracy on a biased sample. This is unglamorous and probably worth more than any architecture change.

4. Split retrieval from ranking.
A single pairwise classifier over the catalog forces a precomputed table. I'd use a two-tower encoder plus approximate nearest neighbor search for candidate generation (cheap, refreshes as the catalog changes, handles new SKUs), then a cross-encoder reranker over the top ~100 that can see the ordered item, the candidate, store context, and the customer together. Retrieval gets coverage; reranking gets precision. It also makes the personalization layer from the last turn much cheaper to apply, since it only touches the shortlist.

5. Hard constraint layer, separate from the model.
Allergens, dietary and religious restrictions, alcohol, pharmacy, infant formula, and large price deltas should be enforced by rules that sit outside the scorer and cannot be overridden by a high probability. The errors that damage trust most are categorical, not marginal — and a learned model trained on aggregate approval will occasionally produce them.

6. Evaluate on the coverage/approval frontier, not a point.
The precision-against-golden-set metric from phase 1 is rank-insensitive and head-biased. I'd report the full curve of approval rate versus coverage as the abstention threshold sweeps, broken out by category and by customer history depth. That prevents shipping a model that looks better only because it helps fewer people.

My notes

What I didn’t understand or am unsure about

Question

Answer

“We compared the Abnormal detection system—a multi-stage pipeline combining behavioral models, downstream rules, and an LLM critic”

How does an LLM critic fit into the Abnormal pipeline? What is its role?

Abnormal says their classification cost is $20/million messages. Based on frontier model cost, LLMs only fire for 0.1% of messages.

This suggests that the critic sits off the default path, invoked on a small minority of messages after cheaper stages have already decided.

The use of the word “critic” could mean that (assuming the prior stages are tuned more for recall than precision) the LLM is being used to suppress false positives but this isn’t stated in the blog post.

A full pipeline could look like

  1. Lighter-weight model/model cascade with a low threshold sends all cases above a certain score to the next step
  2. LLM reviews likely positives based on specified guidelines

“the efficacy comparison reflects this specific evaluation set, not production performance.”

Are there any obvious gaps in the curation of the data set? Beyond the fact that Abnormal’s models are trained to be aligned with Abnormal’s security analyst decisions

  1. Selection — the attack set is conditioned on Abnormal having caught it. Attacks Abnormal never flagged don't reach that queue.
  2. Selection — the safe set is likely hard-negative enriched.
  3. Label taxonomy — "safe" and "graymail" overlap by construction
  4. Temporal — contamination is checked in one direction only
  5. What each system was shown

“returns a decision—attack, spam, graymail, or safe”

How does Abnormal typically deal with these four categories of decisions? What is the action taken based on the decision?

Figure from the post

“Each model ran as a single-pass classifier—given a message, return one verdict (attack, spam, graymail, or safe), with no surrounding pipeline. Each was run under detection-oriented prompting suited to the scenario.”

  1. Are the models given definitions of these verdicts?
  2. Are they given any features besides the message in the email?

(1)

Not stated in the blogpost but they probably were

(2)

Partly stated, and only in the negative. The Executive Summary asserts the models lacked "organizational context—behavioral baselines, sender relationship graphs, and campaign-level signals—that a message-level LLM call cannot replicate." Campaign-level dedup also strips the cross-message repetition that makes commodity attacks trivially detectable.

What's never specified is what "a message" contained: headers, SPF/DKIM/DMARC results, envelope sender vs. display name, URLs, attachments, thread history, or body text alone. That single choice plausibly moves the results more than anything else — authentication headers alone separate a large share of commodity phishing.

2022 Blogpost original

How Instacart Uses Embeddings to Improve Search Relevance

Improve search relevance

My notes

Summary

Context

Searching for items in the search bar on Instacart surfaces merchants and items like in the screenshot below.

Figure from the post

The overall search pipeline consists of multiple components - .

Problem

Improve search relevance

Constraints

Latency, storage(?)

Some key insights/new things I learned

  1. They initialized the two towers in ITEMS with pretrained weights from the Sentence Transformers repository, and fine-tuned the model with Instacart’s in-house datasets
  2. To train model, need positive and negative (query, product) pairs
    1. Positive examples source: Instacart’s search log, customer issued a search query and converted on a product by adding it to their cart
    2. Negative examples source: Given a batch of positive training data, used all off-diagonal (queryᵢ, productⱼ) pairs where i != j
  3. Self-adversarial re-weighting - easy negatives given low weight and hard negatives given high weight
  4. Products converted to embeddings by
    1. Combining product description and metadata alongside special tokens: “”
    2. Passing combined string into Embedding model
  5. Synthetic positives helpful for products with no historical conversions; generated synthetic queries by permuting the various product attributes. Percentage of synthetic data is kept very low to ensure that it doesn’t introduce bias by overwhelming the organic examples
  6. They added auxiliary training tasks to the product tower by using product metadata. Given a product, the model not only learns its representation, but also tries to predict its brand as well as the categories that it belongs to.
  7. Their positive example sourcing method runs the risk that customers add items that are unrelated to the search query but they wanted anyway. To mitigate this, they rank converted products by their conversion frequency for each search query. Then they only keep products above a certain conversion frequency threshold.
  8. Cascade training: first, use a larger, noisier warmup dataset (lower threshold) with shared parameters for the two towers of ITEMS. Then, filtered dataset with separate parameters for the two towers.
  9. Graphs in ITEMS in Action demonstrate that embedding score is more stable and much less prone to popularity bias compared with the raw clickthrough rate; the embedding model provides a fairer baseline for relevance instead of a winner-take-all effect that would be created by using CTR
  10. Serving
    1. Daily offline scoring jobs that pre-compute embeddings for new search queries and new products
    2. Organized into indices using the FAISS library and served in an approximate nearest neighbor (ANN) service with daily updates to the indices
    3. Caching
  11. Evaluation done through A/B testing
    1. +1.2% in mean reciprocal rank (MRR) of the first converted item in search, +4.1% in cart adds per search (CAPS), and a substantial increase in gross merchandise value (GMV)
  12. Semantic Deduplication
    1. Removing semantically identical (similar?) suggestions improves the user experience by presenting to them a set of suggestions that are distinct, diverse, and are more likely to fit their needs

What I didn’t understand or am unsure about

Question

Answer

Using the negative example sourcing they described could still lead to some false negatives if similar queries are included in the search log. Is this a significant enough risk to be concerned about?

This would be pretty rare and neural networks are fairly robust to a small percentage of "noisy labels.”

Standard industry practice is to not treat as negatives other cases where the product ID was identical so this should mitigate the issue somewhat.

The soft loss function also reduces the impact {somehow}.

How exactly does the weighting in the self-adversarial re-weighting influence the training process?

The weights are inside the loss function

Figure from the post

“where the in-batch negative examples are given different weights as a function of the model’s prediction”

How are the weights determined in the self-adversarial re-weighting?

The weights aren’t manually determined. The formula above acts as a scale that assigns a specific weight to every single negative in the batch based on its current score.

How do the special tokens work in product embeddings?

{Full explanation unclear}

It’s able to use these to learn that a match in one category might be more impactful than matches in other categories.

Let’s say the tokens in the [PN] segment were most predictive of the final outcome; the Self-Attention mechanism will then pool more information from the [PN] segment into the final [CLS] vector.

For converting products to embeddings, is there any capacity in their current methodology to incorporate images or reviews?

If not, how could they do it?

Images

  1. Late Fusion
    1. pre-trained Image Encoder (like a ResNet or a Vision Transformer / ViT) to generate a separate "Visual Embedding"
    2. visual vector would be concatenated with the text-based vector generated by the current Transformer
  2. Early Fusion
    1. treat the image as a series of "Visual Tokens" (patches of the image)
    2. These visual tokens would be fed into the Transformer alongside the [PN] and [PBN] tokens. The Self-Attention mechanism would then learn to correlate the word "Milk" in the title with the white jug in the image

Reviews

  1. Aspect Extraction
    1. Use a separate LLM to summarize thousands of reviews into "Aspect Tags" (e.g., "Very Fresh," "Good Value," "Easy to Open")
    2. These tags would be added as new tokens inside the [PAS] (Product Attributes/Stats) segment
  2. Sentiment Weighting
    1. Calculate a "Global Sentiment Score" for the product based on reviews
    2. This score could be used as a scalar multiplier for the final embedding score.
    3. If two products have nearly identical semantic embedding scores, the one with higher positive review sentiment would receive a slight "boost" in the final ranking.
  3. Virtual "Searchable" Reviews
    1. Use the top 5 most helpful reviews as additional text input.
    2. Create a new special token, such as [PRV] (Product Reviews), to append to the product tower input.
    3. This allows the model to match queries that use "consumer language" (e.g., "creamy milk for coffee") which might not appear in the formal manufacturer-provided product name ([PN])

How does their multitasking regime stabilize model training?

User clicks are noisy so adding another objective function mitigates the label noise.

Also helps with Cold Start: when a brand-new milk brand is added to Instacart, it has zero clicks. Without multitasking, the model wouldn't know where to put the new product because it has no behavior data to learn from.

With multitasking, the model sees the brand and category metadata. Because it was trained to cluster those features, it automatically places the new product in the correct neighborhood of the vector space, making it instantly searchable.

How does their multitasking regime ensure that products belonging to the same category or brand are clustered in the vector space?

Embedding vectors sent to brand and category heads. Products with the same brand or category will have their vectors be moved closer together.

How do they determine the weights for the different tasks in the overall loss function?

This isn’t explained in the blogpost.

Industry best-practice would be Grid Search with different weights then look at {evaluation metrics}

Why do they use conversion frequency instead of conversion rate when curating their dataset?

Conversion rate’s utility is impacted by noise from low-incidence products.

Conversion frequency ensures we only use examples where the query and product are very likely to be related

What is meant by “allowing the model to attune to the unique characteristics of the two domains.”

Search queries are usually shorter, informal and more likely to be misspelled. The product strings are cleaner and more organized.

“the few remaining ones not yet cached in FeatureStore are computed on the fly” - how is the caching done?

For the vast majority of the catalog (the "Head" and "Mid-tail" products), embeddings are generated in nightly batch jobs and stored in the FeatureStore.

For each search query coming in, we send a batch of product IDs to the FeatureStore. Any product ID that comes back empty is flagged as a "Cache Miss.". These missing IDs are sent to a dedicated inference service running the Product Tower. The newly computed vector is used for the current search and immediately written back to the FeatureStore so it’s available for the next person who searches for it.

What would be the best way to approach this problem today?

  1. Instacart’s model was unimodal (text-only). Today, you would use a Contrastive Vision-Language (CLIP-style) architecture, like SigLIP 2 or Amazon Nova Multimodal.
  2. LLM-Driven Query Expansion & Rewriting
  3. Two-Stage Retrieval (Dense + Rerank)
    1. Matryoshka Embeddings then Cross-Encoder or ColBERT
  4. Synthetic Data Generation
    1. You feed your product catalog to an LLM and ask it to "Imagine 50 different ways a user might search for this item."

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Learn more about InfoNCE loss, Recall@K, Strong Inductive Bias and Matryoshka Embeddings
  2. Implement local version with synthetic data
2022 Blogpost original

How Abnormal Enhanced Its Detection Platform with BERT Large Language Models (LLMs)

My notes

What I didn’t understand or am unsure about

Question

Answer

Does Abnormal sit in the pre-delivery or post-delivery layer?

Abnormal is a post-delivery architecture. Its API-based architecture connects directly to Microsoft 365 and Google Workspace without MX or DNS changes — the provider delivers the mail, then Abnormal reads it via API and removes anything malicious. Abnormal's own architecture page states the tradeoff: operating post-delivery means messages briefly exist in the mailstore before remediation—typically measured in seconds.

One narrow pre-delivery exception, added Apr 2026: Auto-Forwarding Protection for Microsoft 365 introduces pre-delivery scanning and remediation for emails that are automatically forwarded to third-party tools. That covers one flow, not general inbound.

“where bad actors change a malicious email campaign’s content” - change relative to what? Do they mean that a single campaign can have many different kinds of content, subject etc, potentially using dynamic generation?

Yes, and yes.

“This increases the probability that some of the email variations will bypass traditional email security controls”

What kinds of controls specifically will be bypassed by this?

Content hashes / fuzzy fingerprints — a hash of the body identifies one exact message. Varying "Geeks Squad" vs "Geek_Squad" and randomizing the subscription number changes every hash. Fuzzy fingerprinting tolerates small edits, but the randomized fields are exactly the tokens it keys on.

Subject-line and body regex rules — an analyst-written rule matches a literal string. Rotating between service update and order status falls outside it.

IP and domain reputation, blocklists — reputation accrues per sender IP and domain over time. Fresh domains (geeks-squad51[.]com, geek-squadupdate28[.]com) and rotating IPs have no history to score, so they arrive unclassified rather than bad.

“Abnormal has trained and fine-tuned the BERT models on our own unique data sets”

Are they using BERT embeddings as a feature in a larger model and fine-tuning those embeddings by training on a labeled dataset of final classifications? Or are they training it independently?

Note that ‘trained’ here might mean pre-training on email corpus for domain adaptation.

1. What the sources state. The BERT post says several additional models were built on top of the fine-tuned BERT. "On top" is stacking language. End-to-end training would more naturally be described as one model.

2. What the described architecture requires. The final classifier consumes long-horizon behavioral aggregates — sender-recipient counts, vendor relationships, time since first contact. Aggregations over history aren't differentiable, and a tabular feature set of this shape is almost certainly served to a tree ensemble, which has no gradient path at all. Joint training is therefore not just unattractive but unavailable.

3. What deployment economics favor. A frozen encoder is computed once and cached, serves multiple consumers (campaign similarity, intent scoring, downstream detectors), and doesn't need retraining every time a behavioral feature is added. Joint training would couple a 110M-parameter encoder to a feature set that changes weekly.

Each of these alone is suggestive; (2) is close to dispositive.

Where joint training could still exist: inside the BERT-side models specifically. If an intent classifier is a head sitting directly on the encoder with no behavioral features, that head and the encoder can be trained together. That's joint training within the text subsystem, still separate from the final verdict model. This is a hybrid, not a counterexample.

“BERT can bidirectionally analyze and read text from left to right and right to left. This allows Abnormal to be incredibly accurate and efficient”

What does bi-directionality have to do with efficiency? Are they drawing an implicit contrast with RNNs?

Not super sure but it doesn't matter too much.

“Abnormal is always learning through a combination of supervised and unsupervised machine learning”

What is the role of unsupervised machine learning in Abnormal’s process? Are there any references online?

Inference — three distinct roles, non-overlapping:

  1. Behavioral baselining. The per-customer "history systems" in the ML overview build a distribution of normal communication per user and per relationship. No labels exist for "normal"; the model estimates a density and scores deviations. This is the load-bearing use, and it's what makes the company name literal.
  2. Campaign clustering. Grouping polymorphic variants by embedding proximity requires no labels. The BERT post's claim about determining whether two emails belong to the same campaign is unsupervised at inference time.
  3. Self-supervised pretraining. Continued masked-language-model training on their email corpus adapts BERT's vocabulary to business email. Technically self-supervised rather than unsupervised, though marketing copy rarely distinguishes them.

https://abnormal.ai/learning/anomaly-based-detection

“they exploit trusted relationships between internal and external identities”

Does this mean account compromise or something else?

Path 1: The account is genuinely controlled by the attacker (compromise). Mail comes from the real address with real history. Authentication passes because it is the sender. Covers internal account takeover, which the post names as a target category, and vendor email compromise — a real supplier's mailbox used to send a fraudulent invoice mid-thread.

Path 2: The account is not compromised, only imitated (impersonation). A lookalike domain, a spoofed display name, or a free-mail account with the CEO's name. Nothing is breached; the attacker borrows the identity's credibility. This is where the Geek Squad example in the same post sits — brand impersonation with no compromised asset anywhere.

“It analyzes every email from every identity across thousands of contextual signals”

How can thousands of contextual signals be analyzed within the latency constraints?

First, the premise is looser than it sounds. Post-delivery buys enormous slack. The budget isn't the ~100ms an inline gateway must hit to avoid delaying SMTP; it's seconds, since the message is already in the mailbox and the goal is removal before a click. Abnormal's own architecture page puts remediation at seconds. That's one to two orders of magnitude more headroom than the question assumes.

Then, "thousands of signals" is cheap for the model itself. A gradient-boosted tree scoring a 5,000-dimensional vector is sub-millisecond — traversal depth doesn't scale with feature count, only the features actually used at split points get read. Even a dense neural net over thousands of inputs is a few matrix multiplies. Model inference is not the bottleneck.

The real cost is feature retrieval, and it's mostly precomputed. Inference (not stated in the posts, but standard for this architecture and consistent with their description of layered representations over long time periods):

  • Precomputed aggregates. Sender-recipient history, vendor relationship state, per-user baselines — these are updated in batch or streaming and read as a keyed lookup. Thousands of features collapse into a handful of feature-store fetches.
  • Cheap per-message extraction. Header parsing, URL and attachment metadata, string comparisons. Microseconds each.
  • The genuinely expensive item is BERT. One transformer forward pass dominates everything else combined. Mitigations: batching, caching embeddings for repeated campaign content, distillation or a smaller variant, and cascading — run the encoder only when cheap signals leave the verdict uncertain.

Cascading is the key architectural idea. Not every email pays the full cost. A tiered pipeline scores most mail with cheap features and escalates only ambiguous cases to expensive models. If 95% of traffic resolves at tier one, average latency and cost are dominated by the cheap path while the tail gets full treatment.

What the marketing phrase probably means: thousands of features exist in the feature space, not that thousands of independent computations run per email. Most are precomputed aggregates over shared history, and the count is a proxy for richness rather than per-message work.

“Ability to analyze both North-South and East-West traffic flow”

What does this mean?

North-south — mail crossing the organizational boundary. Inbound from external senders, outbound to external recipients. This is what a gateway sits in the path of.

East-west — mail between identities already inside the environment. Employee to employee, and any lateral movement after an account is compromised.

“Anomaly Detection: Ability to apply the risk models to channels (like email) by aggregating risk signals”

What do they mean by aggregating risk signals?

Three distinct senses it could carry

  1. Combining many signals into one verdict — the ensemble reading. Hundreds of individual indicators (unfamiliar sender, unusual send time, urgent language, first-time vendor) get fused into a single risk score by the model. Most likely primary meaning, since it matches the "thousands of contextual signals" claim in the same section.
  2. Rolling raw events into features over time — the feature-engineering reading. Individual emails get summarized into per-sender and per-relationship statistics (count of prior messages, days since first contact, typical send hours). This matches the first post's description of layered representations built over long time periods, and it's aggregation in the more literal, SQL sense.
  3. Grouping related detections into one incident — the alerting reading. Fifty variants of one campaign surface as a single case rather than fifty alerts.

“automatically remediating attacks”

What does this mean?

The core meaning: removing a delivered message from the mailbox without a human acting. This follows directly from the post-delivery architecture. A gateway blocks — the message never lands. An API-based system can't block on the default inbound path, so its equivalent action is to reach into the mailstore after delivery and pull the message out. "Remediate" is the vocabulary the architecture forces; there's nothing to remediate if you prevented delivery in the first place.

What "automatically" is contrasting with. The manual alternative is an analyst receiving an alert, investigating, running a search across all mailboxes to find every copy of the campaign, and purging each one. One of the customer quotes in the first post names exactly this pain — remediating phishing messages manually consuming hours. Automatic means the platform executes on its own verdict, and the practical claim is that removal happens within seconds of delivery rather than hours.

Actions that plausibly fall under the term, from most to least certain:

  • Deleting or quarantining the message across every recipient mailbox in the campaign, not just the one that triggered detection
  • Moving to a quarantine folder the security team can review, rather than hard deletion
  • Downstream account actions when the trigger is compromise rather than a message — forcing sign-out, revoking sessions, resetting credentials

The first is definitely in scope. The third is a different product surface (account takeover protection) and isn't what this sentence is about.

The tradeoff worth naming: automatic action on a false positive removes legitimate business mail from someone's inbox. That's the operational cost that makes the BERT post's emphasis on reducing false positives load-bearing rather than decorative — a detection system that only alerts can tolerate more noise than one that acts unilaterally.

2023 Blogpost original

Detecting Email Attacks with Generative LLMs

My notes
  1. Combined into a single entry since very closely related
    1. https://abnormal.ai/blog/detecting-email-attacks-with-generative-llms
    2. https://abnormal.ai/blog/how-abnormal-trains-llms

What I didn’t understand or am unsure about

Question

Answer

“Our security analysts now use GPT-4 through the secure Azure API as a tool in their arsenal.”

Who are security analysts? What background do they typically have, and how do they fit into the broader company?

Three distinct functions, MECE by output:

  1. Triage / escalation response — work customer-reported or customer-escalated messages, deciding attack vs. spam vs. safe. Output: a verdict, fast.
  2. Detection quality / labeling — produce the ground-truth labels that train and evaluate models, and audit FP/FN samples. Output: labeled data and error diagnoses. This is the function post 1's labeling section is scaling.
  3. Threat research — characterize new campaigns and attacker TTPs, feed signatures or new detector requirements to engineering. Output: intelligence and detection requirements.

At a detection company, analysts sit between customers and ML engineering, unlike at an enterprise where the SOC is a cost center defending one org. Practical consequence: their labeling throughput is a hard constraint on model iteration speed. That's the actual economic argument in post 1 — GPT-4 doesn't replace analysts, it multiplies the number of messages a fixed analyst headcount can adjudicate, which is why it "helps scale our labeling resources and decrease our response latency for misclassifications".

“Look up suspicious attributes from our vast feature store”

What does it mean to look up suspicious attributes? Why do we need to go to the feature store

I’m not entirely sure. It could mean that they compare the current email’s suspicious attributes to the values of those same attributes for the broader population.

“Summarize an email’s processing logs”

What are email processing logs?

Processing logs are the record of what Abnormal's own detection pipeline did to a message — not mail server delivery logs (SMTP handoffs, bounces), which is the more common meaning of "email logs" elsewhere.

For a single message that would include: which features were computed and their values, which models scored it and what scores came back, which rules or thresholds fired, and the final verdict with the path that led to it. Essentially a trace of one message's journey through the system.

The reason an LLM helps here is volume and format. A single message can generate a large, deeply nested, machine-oriented trace — hundreds of feature values, multiple model outputs, timing data. An analyst asking "why did we let this through?" would otherwise have to read all of it to find the two or three fields that explain the verdict. Summarization turns that into a paragraph. That's also why the next bullet is comparison: the most common diagnostic question is "these two emails look nearly identical, why did we catch one and miss the other?" — which means diffing two traces to find the field that diverged.

“with an internal vector store of labeled messages”

How is this internal vector store of labeled messages being used? Is it being passed in the prompt or some other way?

Not fully clear from the post but one possible interpretation:

The mechanism is retrieval into the prompt. Embed the email under review, run nearest-neighbor search over the store of previously labeled messages, and inject the top matches with their labels as few-shot examples alongside the analyst prompt. GPT-4 then classifies the target with those neighbors as reference points.

“In some instances, we can use GPT-4 at a near-human-trained level of accuracy to determine if a message is malicious.”

A GPT-4 based classification will probably have a few seconds of latency. In what circumstances would this kind of classification be used?

They say “we can reliably scan for messages that our primary system falsely detected or missed” But does that mean that the GPT-4 pipeline is running for every message or just a selected subset?

Answer: a subset, and running after the primary verdict rather than inline.

The source settles the volume question, just in a different section. The model bootstrapping paragraph states that “the costs of running such a large model quickly become infeasible for the billions of messages processed every day“ — that's the stated reason lighter models get trained on GPT-4-generated synthetic data instead. Cost rules out running GPT-4 on everything, and latency rules it out independently.

The stronger clue about when it runs is the phrase "decrease our response latency for misclassifications." You only have a misclassification to respond to after the primary system has already delivered a verdict. So GPT-4 isn't in the blocking path deciding whether to deliver a message — it's auditing decisions already made. Email security has an unusual property that makes this workable: messages can be retracted from a mailbox after delivery. A verdict reversal minutes later still prevents most of the harm, because the window that matters is when the user reads and acts on the email, not when it arrives.

Inference on how the subset gets chosen. The obvious candidates are messages where the primary system's score sits near its decision threshold, since that's where errors concentrate; messages a customer reported; and messages resembling a newly discovered attack, found by searching for neighbors of a confirmed case. Sampling ordinary high-confidence traffic at a low rate would also make sense as a way to measure the false-negative rate you can't otherwise see, but the post doesn't say this.

“for our high-volume detectors”

What do they mean by high volume detector? What would be an example of this kind of detector?

"high volume" describes the traffic the detector sees, not the number of attacks it catches. Post 2 draws the same line explicitly: it says Abnormal traditionally uses lightweight ML models that run on nearly all traffic, and reserves LLMs for specific segments of emails where the added cost is justified. A high-volume detector is one on the "nearly all traffic" side — it scores every message, or a large fraction of them, at a per-message budget of milliseconds and fractions of a cent. Against billions of messages daily, that budget is what forces a small model rather than an LLM.

Example: a BEC / payment-fraud detector. Every inbound message has to be scored for it, because you can't know in advance which emails contain a fraudulent payment request. The base rate is extremely low — a handful of true positives per million messages — so the model runs constantly and fires rarely. That combination is exactly what makes synthetic data attractive: real labeled examples of a specific attack variant are scarce, so GPT-4 generates plausible ones to train the cheap model that does the actual scanning.

The contrast case would be a detector that runs only on messages already flagged as suspicious, or only on those with attachments. Those see a filtered slice of traffic, so a more expensive model per message is affordable.

“we have pipelines to generate synthetic data with GPT-4 that lighter-weight models are trained on”

How do they know that GPT-4 is capable of generating good/useful synthetic data? It doesn't look like they present any evaluation results so it’s not clear to me why data augmentation of this form would be helpful.

The test is downstream, on real data. Train the lightweight model on the synthetic corpus, evaluate it against a held-out set of real emails with human labels. Judging the generations directly — do they look realistic — is a proxy that fails in both directions: plausible emails that teach nothing new, clumsy ones that still fix the right boundary.

The binding constraint is the real evaluation set, not the generation. For a new attack type you might have a dozen confirmed cases. Enough to measure precision, not recall — and at a base rate of a few per million, the false-negative denominator is invisible. The practical move is to ship in shadow mode, where the model scores traffic without acting, and let analysts adjudicate what it surfaces. That measures recall and builds the real labeled set that eventually replaces the synthetic one.

The failure mode to watch is divergence from real attacks. GPT-4's notion of a payroll-diversion email comes from pre-training, not from what attackers send this month. If they've drifted, the pipeline yields a confidently wrong training set and every signal inside the synthetic loop still looks healthy. Only contact with real labeled attacks exposes it.

“Our vast datasets across attack types allow us to train our own large language models”

Do they mean "train" or "fine-tune" here? Training an LLM from scratch seems beyond the capabilities of a company of this size and focus.

Pretty sure they mean fine-tune based on the second post.

“For the task of classifying a given email as a threat, Abnormal traditionally uses lightweight machine learning (ML) models that run on nearly all traffic. However, for certain cases, LLMs have additional benefits that make them worth the cost to run on specific segments of emails.”

How do they determine which emails the LLMs will run on as opposed to the lighter weight models? In other places, they've mentioned that they want millisecond latency, so how does that align with LLM use?

The likely selector is the lightweight model's own score. High-confidence ‘safe’ and high-confidence ‘attack’ get resolved cheaply; the uncertain band in the middle goes to the LLM. That band is small — a fraction of a percent of traffic — which is what makes the cost tolerable, and it's where nearly all errors live, which is what makes it worth paying for. Two other segments plausibly get routed regardless of score: messages matching a newly-launched attack type where no trained detector exists yet, and messages to high-value recipients like finance or executive accounts, where a false negative is expensive enough to justify the spend on every message.

On latency, post 1 resolves the tension. Its labeling section talks about decreasing response latency for misclassifications, which only makes sense after a verdict has been issued — the LLM is auditing decisions, not making them inline. Email tolerates this because messages can be retracted post-delivery, so a reversal seconds or minutes later still lands before the user acts. Post 2 doesn't say whether the LLM classifiers it describes run inline or in that same post-delivery position; my read is post-delivery, for the same cost and latency reasons, but that's inference.

“With larger pre-trained models, we can leverage significantly smaller datasets for newer attacks. In some cases, we can use only a prompt or a single reference email to build a high-precision classifier.”

How did they determine the larger pre-trained models could work reasonably well on smaller data sets? Don't you need a lot of data to make this determination confidently anyway?

Few-shot capability in large pre-trained models was established public knowledge by 2023 — the GPT-3 paper onward. The reasonable read is that Abnormal took it as a known property and tested whether it held on their task, rather than discovering it.

I think they mean that larger pre-trained models need little training data to perform well. They're claiming few labels are needed to build the classifier, not to evaluate it. Those are different pools, and evaluation labels can be produced on demand after the fact.

The specific reason it's cheap here: they claim high precision, not high recall. Precision is measurable by sampling only what the model flagged — run it over a slice of traffic, take the messages it fired on, have analysts adjudicate those few hundred. That's a self-generating evaluation set requiring no pre-existing labeled corpus. Recall would require knowing every attack of that type in the traffic, including the ones nobody found, which is the expensive quantity. The claim is scoped to the half that's cheap to verify, and a high-precision detector that misses cases is still worth shipping when it layers on top of existing ones.

Back-testing covers the rest. Historical traffic is already stored and partly labeled, so a new classifier can be run retroactively against past confirmed attacks of that type — giving a rough recall figure against known cases without collecting anything new.

“Knowing what’s typical gives the large model a powerful “common sense” reasoning ability”

What does this mean?

A lightweight model learns what's typical only from Abnormal's own labeled emails and the features engineered for them. A pre-trained LLM arrives already knowing how business correspondence works — that invoices come from vendors you've transacted with, that a CEO doesn't email a junior employee asking for gift cards, that "urgent wire transfer, don't call me" is a strange thing for a real colleague to write. None of that had to be labeled; it fell out of reading the internet.

Why that reduces both error types. False positives: a message can look statistically odd — new sender, unusual hour, financial language — and still be obviously legitimate to anyone who reads it, like a first invoice from a newly contracted supplier. The LLM can recognize the benign story. False negatives: an attack can be statistically unremarkable on every engineered feature and still be transparently fraudulent in content. The LLM reads the pretext, not just the metadata.

The load-bearing word is "plausible." The model isn't checking facts about the tenant — it's judging whether the situation the email describes hangs together as a normal business interaction. That's a genuinely different signal from behavioral anomaly detection, which is why it composes with it rather than duplicating it. It also fails differently: an attacker who writes a coherent, unremarkable pretext defeats the common-sense check while the behavioral features may still catch the anomaly.

“Using the latest techniques for efficiently fine-tuning these models, like Low-Rank Adaptation (LoRA), we can incorporate our internally labeled email datasets into improving them.”

What do the features and target look like for this kind of fine-tuning?

The input is a text prompt, so everything gets serialized into it: the email itself (headers, subject, body) plus those out-of-band features rendered as text — most plausibly the behavioral history from the feature store, like whether this sender has ever emailed this recipient, prior volume, domain age. The post is explicit that these come from outside the message, which is the interesting part: they're injecting the same signals the lightweight models consume, formatted for a language model to read.

The label is one of attack/spam/safe. Given that post 2 lists explainability as a core reason for using LLMs at all — reasoned justifications, not just "attack" or "safe" — the target is almost certainly label plus written explanation, since a model fine-tuned to emit only a bare label loses that capability. Where the explanation text comes from isn't stated; GPT-4-generated rationales over analyst-labeled emails would be the standard approach.

Let's say fine-tuning is expensive, and I'm not sure yet if it's worth it. How do I determine whether to proceed with it or not? What factors should I consider, and are there smaller-scale evaluations I can use before committing to the full process?

Establish the ceiling before spending anything on training. Run the base model few-shot on a few hundred labeled examples. That number is your floor. Then run the strongest available model — GPT-4-class — on the same set, few-shot. That's roughly your ceiling, since a 7B fine-tune matching it is the good outcome, not the expected one. If the gap between floor and ceiling is small, fine-tuning has little room to buy you anything. If few-shot already clears your production bar, ship that and stop.

Diagnose whether your errors are the kind fine-tuning fixes. Read the failures from the few-shot run. Fine-tuning fixes definitional errors — the model has the reasoning but draws your attack/spam boundary in the wrong place, or won't hold the output format. It doesn't fix missing information: if the model fails because it can't see the sender's history with this recipient, the answer is putting that in the prompt, not adjusting weights. Sorting a hundred errors into those two buckets tells you most of what you need.

Do a scaled-down fine-tune before the real one. LoRA on a few thousand examples for a few hours is cheap relative to the full run. What you're measuring isn't final accuracy — it's slope. Train at 500, 2,000, and 5,000 examples and plot the curve. Still climbing steeply means more data pays; flat by 2,000 means you've extracted what's there and the full run won't help.

The costs that actually decide it are usually not the training run. A fine-tuned model is a hosting commitment, a retraining commitment every time attackers shift, and an eval-set commitment to detect when it's degraded. An API call to a frontier model has none of those. Fine-tuning wins on inference economics at volume — which is Abnormal's stated reason, running on a large share of billions of messages — and on latency. If your volume is low, the arithmetic rarely favors it regardless of accuracy.

How would I convince a business leader that this project was worth doing?

Show Benefit > Cost

  1. Benefit
    1. Reduced False Positives and Reduced False Negatives -> $
    2. Saved Analyst Time -> $
    3. Explainability?
    4. Reduced Infrastructure Cost?
  2. Cost
    1. X hrs of manpower -> $

If you were building this system today, how would you improve it?

The routing economics have changed enough to redraw the diagram. In 2023 the LLM had to be a narrow slice of traffic because inference was expensive and slow. Small fast models are now cheap enough that the "uncertain band" can widen considerably, and more importantly the LLM can move from post-delivery audit into the inline path for a meaningful share of messages. That changes the product, not just the cost line: catching an attack before delivery beats retracting it afterward, because the retraction only works if the user hasn't already acted.

Retrieval and fine-tuning shouldn't be an either/or. Post 2 frames LoRA as replacing the vector store. I'd keep both, because they carry different things. Fine-tuning encodes stable judgment — where your attack/spam/safe boundaries sit, what your output format is. Retrieval carries what changed this week: the campaign that started Tuesday, the vendor domain a customer just onboarded. Baking current threat intelligence into weights means retraining every time attackers shift, which is the wrong update cadence for the fast-moving half of the problem.

The gap I'd treat as most serious is adversarial robustness, which neither post addresses. You have an attacker-controlled text field going into a model that reasons over it. Prompt injection in the email body is the obvious exposure, but the subtler one is that attackers can iterate against your classifier much faster than they could against a feature-based model — send variants, observe which land, keep the ones that do. I'd want injection-resistant prompting where email content is clearly delimited as untrusted data, adversarial examples in the fine-tuning mix, and the LLM deliberately kept as one vote alongside behavioral signals rather than a sole verdict, so defeating it isn't sufficient.

Evaluation should be continuous infrastructure, not a launch gate. The failure mode for a fine-tuned detector isn't a bad launch, it's silent decay as attacks drift away from the training distribution — and recall degradation is exactly the thing you can't see, since you don't count what you never caught. I'd fund a standing labeled stream from analyst adjudication, a held-out set refreshed monthly with recent confirmed attacks, and automatic comparison against a frontier-model baseline on a sample, so the day the cheap model falls behind is a number on a dashboard rather than something you learn from a customer.

2025 Blogpost original

How we built our multi-agent research system

It’s very difficult to predict the required steps for Research in advance so hard-coding a workflow doesn’t make sense. A single agent can struggle with handling both the orchestration and the information gathering/synthesis, both with…

My notes

Summary

Context

People are often interested in exploring complex topics deeply through some kind of research feature in chatbots. The earliest approaches to doing this involved a single LLM using a search tool to find information related to the topic and synthesizing that information into a response.

Problem

  1. It’s very difficult to predict the required steps for Research in advance so hard-coding a workflow doesn’t make sense.
  2. A single agent can struggle with handling both the orchestration and the information gathering/synthesis, both with respect to quality of response and speed.

Solution

Multi-agent system with an orchestrator that can spin up sub-agents for distinct components of the research task.

Some key insights/new things I learned

  1. Citations subagent
    1. {ADD DETAILS}
  2. In our system, the lead agent decomposes queries into subtasks and describes them to subagents. Each subagent needs an objective, an output format, guidance on the tools and sources to use, and clear task boundaries.
  3. https://platform.claude.com/cookbook/patterns-agents-basic-workflows

Design Details

Architecture

  • Multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel
  • [See diagrams in “Architecture overview for Research” section]

Evaluation

Data

Metric(s)

Results

Internal research eval (no details available)

Not shared

90.2% improvement vs single-agent Claude Opus 4

BrowseComp (tests the ability of browsing agents to locate hard-to-find information)

{TO ADD}

Not shared

What I didn’t understand or am unsure about

Question

Answer

“Research work involves open-ended problems where it’s very difficult to predict the required steps in advance. You can’t hardcode a fixed path for exploring complex topics, as the process is inherently dynamic and path-dependent.”

I believe this, but what's an illustrative example that would help me solidify what this looks like a bit better?

Take: "Should we adopt library X for our inference stack?"

A hardcoded pipeline would be: search docs → search benchmarks → summarize. Now watch what actually happens:

  • Step 1 surfaces that X was rewritten in v3, so all benchmarks older than 8 months are irrelevant. The existence of this fork wasn't knowable at query time.
  • Step 2 you now need v3-only benchmarks. You find two, and they disagree by 4×.
  • Step 3 the disagreement traces to one using batch size 1. Now you need your batch size regime — a question that only exists because of what step 2 returned.
  • Step 4 you notice both benchmarks ran on A100s. Now: does the kernel path differ on H100? Another branch spawned by an incidental detail.
  • Step 5 a GitHub issue says the maintainer left in March. Now you're researching project health, which was never in the original decomposition.

“A linear, one-shot pipeline cannot handle these tasks.”

Does one shot here mean a single tool call?

No. It means a fixed, predetermined sequence — however many steps it contains.

  • Linear = no branching. Step 3 is always step 3, regardless of what step 2 returned.
  • One-shot = the plan is committed once, up front, and not revised. Contrast this with the loop the post describes, where the LeadResearcher "synthesizes these results and decides whether more research is needed."

“exploring different aspects of the question simultaneously before condensing the most important tokens for the lead research agent”

How does the sub-agent decide what the most important content to include in the summary is and what the length should be? On the developer's side, is the prompt for the sub-agent just something like “summarize the most important information that you found”?

The lead agent specifies the contract, not the subagent. Principle #2: each subagent needs "an objective, an output format, guidance on the tools and sources to use, and clear task boundaries." So output format is set per-task at spawn time.

The subagent prompt (research_subagent.md, 48 lines) controls compression three ways:

  1. What's important — a four-criterion filter: significant, directly relevant, precise (numbers/dates), from high-quality sources. Judged against its task, not the user's full query.
  2. Length — no word count. A self-set "research budget" in tool calls (under 5 simple → up to 15 multi-part, hard cap 20 calls / ~100 sources), plus "dense in reporting."
  3. Output format — not in this prompt. The lead agent specifies it per task at spawn time.

Two extras: conflicting facts get escalated to the lead rather than resolved, and there's a full section on flagging speculation and low-quality sources.

The full sub-agent prompt can be found here:

https://github.com/anthropics/claude-cookbooks/blob/main/patterns/agents/prompts/research_subagent.md

“Each subagent also provides separation of concerns—distinct tools, prompts, and exploration trajectories.”

I get the different exploration directories part, but why distinct tools? I feel like there are scenarios where two sub-agents are exploring different trajectories, but in the course of those explorations, they require similar tools.

Claude thinks that “Distinct" here doesn't mean disjoint tool sets. distinct trajectories through the same tools are the more common one. Same tool, different call sequence, no shared state.

Unclear what’s actually the case though.

“We found that a multi-agent system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents.”

Is there any reason for the choice of Opus as the lead agent and Sonnet as the sub-agent?

  1. The jobs differ in kind. The lead agent's work is decomposition, delegation, and synthesis — one-shot, high-leverage reasoning where a mistake poisons everything downstream. A bad decomposition can't be recovered by good subagents. Subagent work is search-and-filter: bounded task, clear objective, verifiable output. That's the profile where a smaller model loses least.
  2. Cost scales with the wrong side. There's one lead and N subagents, and the post says subagents do all the substantial information gathering. So subagent tokens dominate the bill. Putting the expensive model on the singleton and the cheap model on the fan-out is where the money is.
  3. Latency is set by the slowest subagent. The post flags this explicitly — synchronous execution means the system blocks on one subagent finishing. Faster subagents shorten the critical path; a faster lead doesn't.

Is there any limit on the number of iterations for the agent system?

Three levels, and they're governed differently.

1. Outer research loop (lead spawns → synthesizes → decides whether to spawn again)
No stated numeric cap, in either the post or the prompts. The post describes it as a decision the LeadResearcher makes each round. Termination is heuristic: stop at diminishing returns, stop if time is running long, then write the report. The lead prompt's example — if you've identified the top 5 startups with high confidence, stop immediately.

2. Subagent count
Soft default 3. Guidelines scale 1 → 2-3 → 3-5 → 5-10. Hard-ish ceiling of 20, phrased as never exceed unless strictly necessary; if you think you need more, restructure.

3. Subagent tool calls
Budget by difficulty: under 5 / ~5 / ~10 / up to 15. Floor of 5 distinct calls. This is the only genuinely hard limit — over 20 tool calls or ~100 sources and the subagent is terminated.

“Multi-agent systems work mainly because they help spend enough tokens to solve the problem.”

Token usage feels like an intermediate causal step, not the actual reason. Why is this framed as token expenditure?

I think their framing is a bit confusing.

1. It's a variance-decomposition claim, not a mechanism claim (source-explicit). The sentence is doing regression-speak: they found token usage explains 80% of performance variance on BrowseComp, with tool calls and model choice bringing it to 95%. "Explains variance" ≠ "is the cause." Tokens are the measured regressor that happens to soak up most of the variation.

2. Tokens are a proxy for the thing you actually can't measure. The real quantity is something like amount of useful serial-plus-parallel reasoning and evidence gathering applied. That has no clean unit. Tokens are its observable shadow — and importantly, the one you can spend deliberately. You can't buy "more reasoning"; you can buy more tokens and hope. That framing is what makes the architecture actionable: the post's next sentence says this validates distributing work across separate context windows to add capacity for parallel reasoning.

“In our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats.”

Is this ceteris paribus or just comparing the average regular chat vs the average agent use?

Most likely observed averages across production traffic, not a controlled comparison.

“Further, some domains that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today.”

  1. Could we not maintain a single global context in addition to each sub-agent's specific context?
  2. Can't dependencies be handled the same way they are in data pipelines or software engineering, where we have a dependency graph?
  1. Shared global context is very token-intensive (N sub-agents x G context length).
    1. Edges are discovered — you learn task C depends on B's output only after B returns. You'd be rescheduling continuously. This isn't a deal breaker, but it would be a little tricky to handle..
    2. The lead agent might not be good at coordinating these dependencies well.

“When a user submits a query, the lead agent analyzes it, develops a strategy, and spawns subagents to explore different aspects simultaneously.”

How does the lead agent determine the decomposition of the different aspects to be explored? On the developer side, is this something we only include in the prompt or are there other ways we can make it better at making this determination?

How it decides (source-explicit, from research_lead_agent.md): classify the query into depth-first (multiple perspectives on one question → 3-5 methodological angles), breadth-first (independent sub-questions → enumerate and split), or straightforward (single investigation). Then plan, with an explicit instruction to define crisp boundaries between sub-topics to prevent overlap.

Beyond the prompt — five levers, roughly in order of payoff (inferred):

  1. Evals on decomposition itself. Grade the plan, not just the final report. Cheap, fast, and it's where errors originate.
  2. Fixed decompositions for recurring query shapes. If "compare N vendors on M criteria" is common, template it rather than re-deriving.
  3. A cheap scoping search before planning. Decomposing "compare EU tax systems" without first knowing the member list produces bad boundaries. The lead prompt does this — retrieve the list, then split.
  4. Replanning after the first round. The post's synthesize-and-decide loop is a second chance; a strong first decomposition matters less if gaps get caught.
  5. Model choice. Their Opus-lead setup is itself a lever here.

“In contrast, our architecture uses a multi-step search that dynamically finds relevant information, adapts to new findings, and analyzes results to formulate high-quality answers.”

This is a little vague. Let's flesh this out with some details and examples.

"Adapts to new findings" is the real differentiator, and it's implemented as an explicit OODA loop the subagent runs after every tool result — observe what's been gathered, reorient on what's still missing, decide on the next tool call, act. The lead prompt's own example shows why this can't be precomputed: for "compare EU country tax systems," first spawn a subagent to retrieve the current member list, then reason about which metrics are worth comparing, then split subagents by region. Step two's content is a function of step one's output, so no fixed plan could have specified it.

"Analyzes results" turns out to mean source skepticism specifically, not generic reasoning. The subagent is told to catch pages that speculate in future tense ("could," "may," financial projections) and report them as predictions rather than as events that happened, to prefer original sources over aggregators, and to escalate conflicts it can't reconcile to the lead agent rather than silently picking a winner.

The other two clauses are smaller. "Multi-step" is a 5–15 tool-call budget with a search-then-fetch-the-full-page loop. "Dynamically finds" means queries are written per step under a five-word guideline, broadening when hits are thin.

“evaluates tool results using interleaved thinking”

Elaborate on how this works and why it’s used.

How it works. Standard extended thinking happens once, before the model's first output. Interleaved thinking lets the model emit a thinking block after each tool result and before the next tool call, so reasoning is inserted at every step of the loop rather than only at the front. In the Claude API this is a beta feature, and the thinking blocks stay in the conversation history so later steps can see the earlier reasoning. (Mechanism is from the docs the post links to, not the post itself.)

Why is it used. Without it, the mapping from tool result to next tool call is close to reflexive — the model sees ten search snippets and immediately fires the next query. The post names three things that the reasoning step is supposed to do: evaluate the quality of what came back, identify what's still missing, and refine the next query accordingly.

Where you can see it load-bearing. Several instructions in the subagent prompt are only enforceable if there's a reasoning step between result and next action:

  • The research budget — deciding you've hit diminishing returns requires assessing the results you have.
  • The broaden-or-narrow rule — you can't tell whether to widen a query without judging whether hits were thin or abundant.
  • The source-skepticism checks — spotting that a page is speculating in future tense, or is an aggregator rather than an original, is an evaluation of content, not a lookup.

“The LeadResearcher synthesizes these results and decides whether more research is needed.”

How does the lead agent determine whether more research is needed? On the developer side, is this something we only include in the prompt or are there other ways we can make it better at making this determination?

How it decides: it compares the facts gathered so far against the plan it wrote at the start, and stops at diminishing returns — the prompt's example is that if you've been asked for the top 5 startups and have five with high confidence, stop immediately rather than confirming further.

Best lever beyond the prompt: have the lead emit its plan as an explicit checklist of sub-questions. Then sufficiency is "which items are unanswered" rather than a judgment call over a pile of text. Second-best: have subagents report what they couldn't find, so the lead can distinguish "no such fact exists" from "nobody looked."

“CitationAgent, which processes the documents and research report to identify specific locations for citations”

How does the CitationAgent find citations?

Mechanism (from citations_agent.md in the same cookbook directory): it's a post-hoc mapping step, not a retrieval step. It receives the already-written report plus the source documents the subagents gathered, and its job is to attach claims to the documents that support them. It is explicitly forbidden from changing the report text — it only inserts citation markers.

That constraint is the whole design. The report gets written first with no citations at all — the lead agent's prompt tells it never to include markdown citations or a sources list, because a separate agent handles that. Separating the two means the writing model can't quietly reshape a claim to fit a source it half-remembers.

The limitation worth naming (inferred): this catches claims that don't match any gathered source, but it can't catch a claim that was faithfully sourced from a bad source. Attribution accuracy and factual accuracy are different things — which is presumably why their eval rubric scores them as separate criteria.

“when given a flawed MCP tool, it attempts to use the tool and then rewrites the tool description to avoid failures”

What happens when the original MCP tool structure and description change? Is the agent checking the original version every time to make sure it hasn’t changed?

Not addressed in the post at all. It describes the tool-testing agent as a one-off improvement process producing a better description, and reports the 40% reduction in task completion time. Nothing about re-validation, versioning, or drift.

Inferred, and I think this is a real gap: the rewritten description is a cached artifact derived from behavior observed at one point in time. When the upstream tool changes, that cache goes stale silently — and the failure is worse than having no rewrite, because agents now confidently follow a description that no longer matches reality.

Checking the original on every call is the wrong fix — it's a per-call cost for a rarely-changing input. The standard approach is to hash the upstream schema and description, store the hash alongside the rewrite, and re-run the testing agent when the hash changes. That's cheap and catches declared changes.

What it doesn't catch, and this is the harder half: behavioral drift with no schema change. The description and signature stay identical, but the endpoint now returns paginated results, or rate-limits differently. Only re-running the tests catches that, which argues for periodic re-validation rather than change-triggered alone.

“the subagents use 3+ tools in parallel”

Wouldn’t this result in the same issues that necessitated using subagents in the first place? Coordination, dependencies, etc.

I'm not super sure, and I didn't find Claude's answer (attached below) very convincing.

No — because these are parallel tool calls, not parallel agents. One subagent issues three calls in a single turn, gets all three results back into its own single context, and reasons over them together. There's exactly one decision-maker and one context window, so there's nothing to coordinate. Contrast with subagents, which each have separate contexts and can't see each other's work — that's where coordination cost comes from.

It looks like a lot of their improvements came from ideating and experimenting with different prompts in various parts of their pipeline. What’s the best way in general to do this kind of experimentation?

1. Observe — build visibility before touching anything
Simulate the system with the exact production prompts and tools, and read transcripts step-by-step. This is what surfaced their failure modes: agents continuing past sufficient results, writing overly verbose queries, selecting wrong tools. You can't edit a prompt usefully without an accurate model of what the agent is currently doing.

2. Change — generate the candidate fix
Two sources. Your own read of the transcript, and the model's: give Claude a prompt plus a specific failure trace and ask it to diagnose. They found Claude 4 models good at this, and built a tool-testing agent on the same idea.

3. Measure — decide whether it worked
Start with ~20 test cases, not hundreds; early changes move success rates by tens of points and are visible with few examples. Grade with a single LLM-judge call producing a 0.0–1.0 score plus pass/fail — they found multiple specialized judges less consistent. Keep human testers permanently, since they catch what rubrics can't: their testers found the agents preferring SEO-optimized content farms over academic sources, in outputs that otherwise looked fine.

“We also proactively mitigated unintended side effects by setting explicit guardrails to prevent the agents from spiraling out of control.”

What does this mean?

Resource caps. Subagents: hard limit of 20 tool calls and ~100 sources, or the subagent is terminated. Lead: never more than 20 subagents, with the instruction that needing more means you should restructure instead.

Termination rules. Stop at diminishing returns; if the process has run long, skip further subagents and write the report immediately.

Scope boundaries. Each subagent gets explicit limits to prevent research drift, and the lead is told to define crisp boundaries between sub-topics so agents don't duplicate each other.

Harm constraints. The lead is forbidden from spawning subagents to research anything promoting hate speech, violence, discrimination, or catastrophic harm, and must add constraints when a query is sensitive.

What "spiraling" refers to (source-explicit): the failures named earlier in the post — spawning 50 subagents for a simple query, searching endlessly for sources that don't exist, agents distracting each other with excessive updates.

“Traditional evaluations often assume that the AI follows the same steps each time”

In what sense specifically do traditional evaluations assume this?

What breaks with agents isn't nondeterminism per se — it's that the correct output is no longer unique. Two subagents can search three sources versus ten, use different tools, and both produce good research reports with different wording, different citations, different structure. Exact-match scoring calls one of them wrong. That's why the post moves to LLM-as-judge: it needs a grader that scores quality rather than identity.

What’s the right way to think about overfitting in the context of evaluating and improving agentic systems, especially with very limited sample sizes like they have?

The specific failure mode to watch for: prompt changes that encode your test cases. "Prefer academic PDFs over content farms" is a general heuristic. "When researching semiconductors, check TSMC investor relations" is a memorized answer to one eval item that will not transfer. Their own principle — instill heuristics rather than rigid rules — is partly an anti-overfitting stance, though the post doesn't frame it that way.

Run each case multiple times so you can separate prompt effect from run-to-run variance. Grow the suite as scores rise — the sample size you need scales inversely with the effect size you're trying to detect.

“We used an LLM judge that evaluated each output against criteria in a rubric”

  1. Which LLM did they use and why?
  2. How are each of the criteria in the rubric determined?
  3. What does the full pipeline for this look like?
  4. “found that a single LLM call … was the most consistent and aligned with human judgements”. I can understand token savings but why better performance through a holistic call?

1. Which model — not stated. The post never names the judge model. Inferred: most likely the strongest model available to them at the time, since judge quality is the ceiling on eval quality and the eval runs offline where cost matters less than in production. Using the same model family as the system under test is the standard critique here — self-preference bias is well documented — but the post doesn't address it either way.

2. Rubric criteria — the criteria are stated; the derivation isn't. They are factual accuracy, citation accuracy, completeness, source quality, and tool efficiency. Inferred derivation: these read as one criterion per observed failure mode. Source quality maps directly to the content-farm bias human testers found; tool efficiency maps to the over-investment problem their scaling rules were written to fix. That's the general method — watch it fail, name the failure, add a criterion.

3. Pipeline (partly stated, partly inferred). Stated: ~20 queries representing real usage; each output scored by one judge call returning a 0.0–1.0 score plus pass/fail; used to scale evaluation across hundreds of outputs; run alongside human testing. Inferred but not described: how scores aggregate across the suite, whether cases are re-run for variance, and how the judge itself was validated against human ratings.

4. Why one holistic call beats several specialized ones — the interesting question, and the post only reports the finding. Three mechanisms, all inference:

The criteria aren't independent. A report that omits a key aspect is usually also weaker on source quality — it didn't look hard enough. A judge seeing everything at once can weigh a completeness gap against how thoroughly the rest was researched. Separate judges each see a slice and can't make that trade, so their scores get combined by an aggregation rule you invented rather than by judgment.

You have to hand-pick the weights. Five separate scores need a combination function. One call lets the model do the weighting implicitly, which apparently matches humans better — plausible, since humans also grade holistically rather than summing subscores.

More calls, more variance. Five noisy judgments combined can be noisier than one, especially if the aggregation is non-linear or if any single criterion can veto.

Worth noting: their finding says consistent and aligned with human judgements — those are two different claims. Consistency is easy to explain by variance. Alignment is the one that suggests holistic grading genuinely mirrors how people evaluate research.

“In our case, human testers …”

What are human testers asked to do specifically and do they require domain expertise or can they just be randos?

Inferred: the failures they list split cleanly by what kind of person can catch them.

Anyone can catch system failures and unusual-query hallucinations, because you don't need to know the subject to notice the agent broke or answered a weird question with obvious nonsense. That's exploratory testing — hand people the system, let them throw odd queries at it, report what looks wrong.

Only a domain expert catches the source-quality bias. Recognizing that a report leaned on content farms instead of the authoritative literature requires knowing what the authoritative literature is for that topic. Nothing in the output looks wrong to a non-expert — it reads fluently and cites real pages.

So the practical answer is both, doing different jobs. Non-experts for breadth: many people, many strange queries, catching crashes and obvious failures. Experts for depth: few people, queries in their own field, catching the plausible-but-wrong. If you can only staff one, the expert is more valuable, because the failures they catch are exactly the ones your automated rubric and your non-expert testers both score as passing.

“Adding source quality heuristics to our prompts helped resolve this issue”

What is a source quality heuristic?

The post doesn't define the term, but research_subagent.md contains the heuristics themselves — so here's what they actually look like.

A source quality heuristic is a text-level cue that predicts unreliability, stated concretely enough that the model can check for it without knowing the subject matter. Their prompt names several:

Speculation dressed as fact — future-tense narrative, "could" or "may," financial projections, quoted superlatives. The instruction is to report these as predictions rather than as things that happened.

Provenance problems — aggregators rather than the original source, false authority, passive voice paired with nameless sources, unconfirmed reports.

Language that signals an agenda — marketing copy for a product, spin, general qualifiers without specifics, cherry-picked data.

Why heuristics rather than a source whitelist: you can't enumerate good domains in advance for arbitrary research topics, and the failure they were fixing was rank-based — content farms outrank academic PDFs and personal blogs, so anything keyed to search position reproduces the bias. These cues are checkable on any page.

The connected instruction that makes them work: the lead agent is told to define what constitutes a reliable source for that specific task and to list unreliable sources to avoid — so generic heuristics get supplemented with per-task guidance.

“letting the agent know when a tool is failing”

What does it mean to let an agent know when a tool is failing? How do you do it/what do you say?

What it means concretely: when a tool call fails, you return the failure as a tool result the agent reads rather than raising an exception that kills the run or silently retrying behind its back. The agent then treats the failure as evidence and adapts — different tool, different query, or noting the gap in its report.

What you actually put in the result. The useful content is whatever tells the agent what to do differently:

  • What failed and how — timeout, rate limit, 404, empty result set, malformed response. These imply different next moves: a rate limit says wait or switch tools; a 404 says that URL is dead and try another source; an empty result set says the query was too narrow.
  • Whether retrying could help, since "transient, try again" and "this will always fail" lead to opposite actions.
  • What alternatives exist, if any.

What to avoid: raw stack traces, which burn tokens and don't tell the model what to do; and bare "Error", which gives it nothing to reason about, so it usually retries the identical call.

Why this works well (the post's phrasing is that it works surprisingly well): the agent already has a reasoning step after every tool result via interleaved thinking. A failure message lands in exactly the slot where the model is deciding what to do next, so error handling gets to reuse machinery that's already there — no separate recovery logic.

The complement, from the post: this pairs with deterministic safeguards — retry logic and checkpoints — so the model handles the judgment calls while infrastructure handles the mechanical retries.

“Agents make dynamic decisions and are non-deterministic between runs, even with identical prompts”

Is there no way to make them deterministic? I know setting the temperature to zero should ensure identical LLM call output for identical prompts, but is the non-determinism coming from the tool calls?

  • Source 1 — the environment. Tool results differ between runs. Search rankings shift, pages update, results vary by geography and personalization. Even a perfectly deterministic model gets different inputs and therefore takes different actions. This is the dominant term for a research agent.
  • Source 2 — the model. Temperature 0 removes sampling noise but not numerical noise. Floating-point addition isn't associative, and GPU reduction order varies with batch composition and scheduling, so two near-equal logits can flip which is the argmax. Your request is batched with other traffic, so this varies run to run. One flipped token early changes the entire downstream trajectory.
  • Response to source 1: record-and-replay. Log every tool call and result, replay against the cache. This fully pins the environment for debugging.
  • Response to source 2: nothing available to you as an API consumer.

“Evaluating agents that modify persistent state across multi-turn conversations presents unique challenges”

What do they mean by persistent state over here?

What it means: state outside the conversation that survives after the turn ends, and that the agent's own actions change. The distinguishing property is given in the next sentence — each action changes the environment for subsequent steps, creating dependencies. Contrast with the Research feature itself, which the appendix explicitly calls read-only: web searches don't alter the web, so every step starts from the same world.

Concretely, in their tool inventory: the Google Workspace and integration tools the post mentions — sending an email, creating a calendar event, editing a Drive doc, updating an Asana task. Also the memory store the lead agent writes its plan to, and the filesystem artifacts the appendix describes subagents creating.

Why this specifically breaks evaluation. In read-only research, you can re-run the agent freely and score the final report. Once the agent mutates state, step 3 sees a world step 2 changed, so you can't score turns independently or replay cheaply — and an action taken and then undone leaves the same end state as an action never taken, but they aren't equivalent. Their answer is to score the final state rather than the process, with checkpoints for long workflows.

“We implemented patterns where agents summarize completed work phases and store essential information in external memory before proceeding to new tasks”

  1. How is the length and content of the summary determined?
  2. How do agents know when to access the memory?

Neither question is answered in the post. The appendix describes the pattern in one paragraph and gives no prompt text, no length rule, no retrieval trigger. The published cookbook prompts don't cover it either — they're for the read-only research loop, not the long-horizon memory pattern. So everything below is inference.

1. Length and content. The reliable signal in the post is what gets retrieved: it names the research plan specifically as the thing recovered from memory rather than lost at the context limit. That points to a fixed, small set of durable items — plan, decisions made, constraints discovered, open questions — rather than a proportional summary of everything that happened.

The design principle I'd apply: summarize by schema, not by ratio. "Compress the last 50k tokens to 5k" degrades unpredictably, because the model has to guess what matters. "Fill in these five fields" is stable across phases and makes omissions visible. Length then falls out of how much actually belongs in each field.

2. When to access it. Two triggers, and they're structurally different:

On spawn — the appendix says fresh subagents are created with clean contexts and continuity is maintained through careful handoffs. That's not the agent deciding to read memory; it's the harness injecting the relevant state into the new context at construction. Deterministic, no judgment involved.

On approaching the limit — the agent retrieves stored context rather than losing previous work. This one has to be prompted, since the model needs to know memory exists and that reading it is the move when context is running short.

The general principle: make retrieval deterministic wherever you can. An agent that must decide to check memory will sometimes not, and the failure is silent — it proceeds confidently without information it had.

“maintaining continuity through careful handoffs”

What do handoffs mean here?

A handoff is the packet of state you write into a fresh agent's context at the moment you create it. The setup in that passage: context is nearly full, so instead of continuing in the exhausted context you spawn a new agent with a clean one. The new agent inherits nothing automatically — a new context starts empty. So whatever it needs to carry the work forward has to be explicitly assembled and passed in. That assembly is the handoff.

"Careful" is the operative word, because the handoff is a lossy bottleneck. Everything not included is gone. Include too little and the new agent repeats work or contradicts decisions the old one made; include too much and you've refilled the context you just cleared.

What's actually in one, inferred: the plan (the post names this as the thing retrieved from memory), decisions already made and why, constraints discovered, what's been completed, what's outstanding, and pointers to stored artifacts rather than the artifacts themselves.

The connection worth noticing: this is the same problem as a subagent returning findings to the lead — compress a full context into something small enough to transmit. The appendix's filesystem-artifact pattern is the fix for both: pass lightweight references, let the recipient fetch what it needs.

If I'm working at a company that could benefit from using some kind of multi-agent system such as the one outlined in this blog post, is it better to build our own system or can we use Anthropic’s system through an API? If both options are available, what factors should be considered in making this determination?

Build vs. buy, resolved: there are three tiers, not two options.

Messages API — you write the agent loop yourself. This is what the blog post's team did, and it's the most work.

Claude Agent SDK — you get the loop; you write the orchestration. Subagents are a built-in primitive with per-agent tool allowlists and model overrides, and several guardrails the post describes hand-rolling are now just config: spawn depth, concurrency cap, dollar spend limit. Runs on your infrastructure.

Claude Managed Agents — Anthropic hosts the harness, state, and sandbox. Public beta. Strongest for single long-running agents; its multi-agent orchestration is still research preview.

How to choose, in the order the questions actually bind:

  1. Compliance first. Managed Agents isn't eligible for Zero Data Retention or a HIPAA BAA, since state lives server-side. In a regulated domain this eliminates the hosted option before anything else is considered.
  2. Does your problem fit multi-agent at all? The post is explicit: heavy parallelization, information exceeding one context window, many complex tools. Poor fit where agents must share context or have tight dependencies. This is the most common way these projects fail.
  3. Is value per task high enough? At ~15× chat token usage, a research report that saves an analyst a day clears the bar; a ticket classifier doesn't.
  4. Can you staff the eval loop? The post's central message is that the gap between prototype and production is prompt iteration, evals, tracing, and human testing. If you can't fund that ongoing, buy something maintained even at worse fit.

How would you improve this system if you were building it today?

1. Fix the information bottleneck between agents. The lead only ever sees each subagent's final report — a lossy summary of a full context. Two consequences: silent absences (the lead can't tell "no such fact exists" from "nobody looked"), and information loss the post itself names as the game of telephone. The fix is in their own appendix but not in the shipped prompts: subagents write full findings to a filesystem and return lightweight references, plus an explicit "could not find X" field. This is the highest-leverage change because everything downstream — the sufficiency check, the final report, the citations — is bottlenecked on what the lead can see.

2. Make the sufficiency check mechanical rather than judgmental. Today the lead reasons about whether its facts are enough, which is a vibe check over a pile of text and has no budget ceiling — the failure that produced 50 subagents for simple queries. Have the lead emit its decomposition as an explicit checklist at planning time, then "is more research needed" becomes "which items are unanswered." Add an orchestrator-level tool-call and spend cap, which the leaf level already has and the coordinator doesn't. (The Agent SDK now ships depth, concurrency, and maxBudgetUsd limits for exactly this.)

3. Go asynchronous. The post flags this itself: synchronous execution means the whole system blocks on the slowest subagent, the lead can't steer mid-flight, and subagents can't hand off to each other. Given that dependencies are discovered rather than declared, the value is less about latency than about acting on a blocking result the moment it lands instead of at the end of the round. Highest complexity of the three — state consistency and error propagation get materially harder — which is why I'd rank it third despite the ceiling being highest.

2026 Blogpost original

Battling Scams at Scale: Inside Doppel’s High-Throughput ML Platform

My notes

What I didn’t understand or am unsure about

Question

Answer

The blog post draws a distinction between detection and monitoring. Why do they draw this distinction, and what would it look like in practice?

Detection tries to find new malicious content whereas monitoring watches already-actioned entities. Caching, featurization, prediction priors, crawl cadence would all look different here.

“Manually orchestrating each step not only slowed development …”

What does it mean to manually orchestrate each step?

  1. SQL in a notebook with results dumped to CSV or ad-hoc table
  2. Feature logic lives in both training notebook and serving service
  3. Training done in notebook; hyperparameters and metrics not stored in a structured way
  4. Bespoke Dockerfile per model
  5. Deployment done through hand-edited config and manual endpoint creation

“Our earliest models were trained on a combination of … ”

  1. Where could they get data for confirmed takedowns from real impersonation cases?
  2. Are the heuristics used as labels or features? If they’re used as labels, shouldn’t we account for label noise?
  1. From their own takedowns and customer-reported cases. Final labels can be confirmed by registrar/host/platform response or a post-takedown crawl.
  2. I think both but not fully sure. If they’re used for ranking for human review, the negative impacts of label noise can be mitigated.

“prioritized models that could help rank content for human review, not make binary decisions”

How would a model that's focused on ranking content for human review differ from one that's designed to make decisions itself in terms of architecture, objective function etc?

The autonomous decision model

  1. could use an abstention head or conformal prediction to route low-confidence cases to humans.
  2. Requires calibration (perhaps using Platt scaling or isotonic regression).

The ranking model

  1. Uses slightly different metrics like precision@k, NDCG, mean average precision where k = daily SOC review capacity.

“Our initial labeled sets risked overfitting to high-confidence edge cases”

Why is this true?

When positives are extreme, cheap correlates separate them perfectly. A blatant phishing domain may have high string entropy, a brand token, a recent registration, and a suspicious TLD all at once. Gradient descent has no reason to learn anything subtler — the shortcut achieves near-zero training loss.

“Our serving infrastructure must process a continuous stream of 100 million+ URL checks per day, maintaining sub-100 ms P99 latency under bursty traffic”

How is latency measured in this case? Is it the time taken between the URL coming online and it being detected, or something else?

No, it’s request-service latency - interval from a query arriving at the feature platform to a score being returned.

“Burdened by tuning constraints, driving up false positives and missing real scams”

What do they mean by tuning constraints, and why would that lead to this impact?

You can’t modify the objective to account for asymmetric costs of FN vs. FP.

If the vendor only returns final classifications and not scores, you can’t tune the threshold for different customers (who may have different base rates).

Even if they do return a score, that score may not be calibrated appropriately.

“Low-latency: supporting both bulk-style and point-based feature generation and inference in real time“

What's the difference between bulk style and point-based?

A single entity vs many entities.

Training is typically done in a batch, whereas serving is point-based.

“The features we define are “resolved” by functions”

What does it mean for a feature to be resolved by a function?

A feature is just a declared name and type (Url.phishing_probability: float) holding no value; a resolver is a function registered as the thing that can produce it, and "resolving" is running that function on demand. Because the resolver's type signature declares its inputs and outputs, Chalk derives the dependency DAG automatically rather than you wiring it — which is what makes caching transparent, lets batch and real-time share the same definitions, and turns a model call into just another node.

“This lets us treat model scores just like any other resolver, composing them seamlessly into downstream workflows.”

What does this mean?

A model's output is written as a feature like any other, so nothing downstream can tell whether Url.phishing_probability came from a neural net or a string-entropy calculation. Both are just nodes in the same DAG.

“has_suspicious_language: A boolean flag for whether the HTML content language is in a suspicious language list”

What do they mean by whether the HTML content language is in a suspicious language list?

This is confusingly worded I'm not sure. The literal reading would imply suspicious languages (e.g. Urdu) are a signal for violative content but that seems overly punitive.

“Behind the scenes, we package each model and its dependencies into custom Docker containers and serve them via lightweight Cloud Run services.”

  1. Why are the models packaged into custom Docker containers
  2. What are Cloud Run services and what’s their role here?

1. Custom Docker containers: Each model gets its own image so its dependencies, runtime, and weights are isolated and pinned — six models can't be forced onto a lowest-common-denominator environment, and a version becomes an image tag you can roll back. The stated payoff is modularity, testability, and independent scaling/versioning/monitoring per model.

2. Cloud Run: Google's serverless container platform — you hand it an image and it runs instances behind an HTTPS endpoint, autoscaling with request volume (including to zero) with no cluster to manage. Here, each model is one Cloud Run service that Chalk's resolver calls over HTTP, so Chalk owns the dependency graph and Cloud Run owns the compute — though scale-to-zero conflicts with their sub-100ms P99, so hot-path models almost certainly pin minimum instances.

“Raw input caching: Store crawled HTML in the online feature store with a TTL to eliminate redundant fetches and parsing.”

What’s the alternative to storing them in the online feature store? Where would they be stored if we didn’t cache?

Without caching, it isn't stored at all. The crawler fetches the page, features get computed from it in-memory, and the HTML is discarded when the request ends. The next check on that URL re-fetches over the network and re-parses from scratch. "Where would it live" presumes persistence that doesn't exist by default.

“Entity-level feature cache: Precompute and cache feature vectors for features related to customers to slash per-request computation.“

Is the idea here that we’re making multiple predictions per customer (e.g. against different accounts) so we want to store customer-specific features instead of recomputing?

Yes.

What sits in it (inferred; the post only says "features related to customers"): the brand's official domain list, canonical name and token variants, logo embeddings, expected locales, historical infrastructure fingerprints for their legitimate properties. Note these are also the reference side of the comparison features they mention elsewhere — brand-name keyword similarity, page-content-vs-claimed-brand inconsistency. Those are two-argument: one side varies per URL, the other is fixed per customer. Only the fixed side is cacheable.

Why it's cheap to cache and expensive to recompute (inferred): customer entities change on the order of days; URLs arrive at ~1,150/sec. Recomputing a logo embedding per request would be absurd. And the cardinality is tiny — hundreds or thousands of customers against 100M+ daily URL checks — so this is the highest hit-rate layer of the three by a wide margin.

“Cache model outputs for frequently seen URLs”

Is the idea that the same URL is appearing multiple times for the same brand across different days or for different brands on the same day? If it’s the former, wouldn’t caching the model output mean that updates to the page wouldn’t be detected?

I think both situations could apply. If it is indeed the former, I think that this is a gap in their methodology. Instead of caching the URL on its own, you could do a URL-content cache.

“Leverage Chalk’s query planning DAG to reconstruct the exact computation graph for every score”

What does it mean to reconstruct the graph?

{TO ADD}

“Centralized label management with dbt”

How does dbt let you manage labels?

{TO ADD}

“sharing the same code paths for batch and real-time compute eliminated drift and made model behavior predictable.”

Sharing the same codebase seems like the default. Is there any reason why the company would have done it differently before realizing it causes drift?

{TO ADD}

“combined with custom Docker containers on Cloud Run smashed redundant work”

What do Docker containers and Cloud Run have to do with minimizing redundancies?

{TO ADD}

“We baked Pydantic schema checks, dataset-drift alerts, and “shadow” inference tests into our CI/CD pipelines”

  1. What would dataset-drift alerts look like for this use-case?
  2. What do they mean by shadow inference tests?
  1. Examples of drift for their use-case:

Figure from the post

What the alerts likely are (inferred): per-feature distribution tests (PSI, KL divergence, KS) comparing a rolling window against the training reference; null-rate and cardinality monitors catching upstream crawler breakage; score-distribution shift on the output side, which is the cheap early-warning signal since it needs no labels.

  1. Shadow inference tests: running a new model version against live production traffic while its predictions go nowhere — logged and compared, never acted on. The current model continues serving.

What it catches that offline eval can't (inferred):

Check

Why offline eval misses it

Training-serving skew

Features computed by the live resolver path may differ from the training path — you only see it on real traffic

Score-distribution shift

New model's scores may sit at a different scale, silently invalidating the deployed threshold

Disagreement analysis

Where new and old disagree is a directly reviewable sample, and a cheap labeling queue

Latency and resource profile

Real P99 under real burstiness, not benchmark conditions

2019 Blogpost original

Model Calibration when Subsampling Negatives

My notes

What I didn’t understand or am unsure about

Question

Answer

“Because of this class imbalance we want to include every single positive example (attacks) but only a portion of negative examples (safe emails).”

What would happen if we included all examples?

  1. Increased cost/latency
  2. Correct calibration

“When we train a model on this data we may get good predictive power, but those predictions will not be actual probabilities (most obviously, the mean is shifted, your model will predict the average class probability to be 0.001 when in reality it is 0.0001)”

Does this depend on which model we use or will that happen regardless?

Regardless of model, for any model trained to minimize a proper scoring rule — log loss, Brier score. That covers logistic regression, gradient-boosted trees, and neural nets with a sigmoid or softmax output. These objectives are minimized when the predicted probability equals the empirical conditional probability in the training distribution. Subsampling changes that distribution, so the optimum shifts with it. It's a property of the objective, not the architecture.

“Matching the distribution of a model’s prediction with the real distribution: pred(class=1| example) ≈ prob(class=1| example)”

This formula needs to hold on average, right, not per prediction the way they’ve written?

What the post writes is strong (or perfect) calibration: the prediction equals the true conditional probability for every individual example. That's essentially the definition of the Bayes-optimal predictor. It's a coherent quantity if you assume labels are stochastic given x, but it's unattainable and unmeasurable — you observe one binary outcome per example, never its underlying probability.

What calibration actually means operationally is a conditional average over the prediction bucket:

E[Y | pred(x) = p] = p, for all p

Among all emails scored 0.3, roughly 30% should be attacks. This is verifiable — bucket predictions, compare to empirical rates — and it's what a reliability diagram plots.

Why the weaker version is the right target. It's also what the post's own machinery delivers: isotonic regression fits a monotone map from raw score to empirical rate within score bins. That can only correct bucket averages; it has no per-example information and cannot recover a per-example truth.

“In the example above, if we were to choose a threshold of 0.5, it will achieve 96.6% precision on the subsampled data, but only 74.2% precision on the full dataset.”

Why did the precision go down on the full dataset?

Because you added back the 90% of negatives you had thrown away, and none of the positives changed.

Read it off the left panel at τ = 0.5. The positives curve (orange) is identical in both settings — every positive was kept. The dashed navy line is subsampled negatives above threshold; the solid blue is full negatives above threshold, sitting exactly 10× higher everywhere, because the subsample kept 1 in 10.

So precision = TP / (TP + FP) has a fixed numerator and a denominator whose false-positive term multiplies by 10 when you restore the full negative population.

Above the threshold, that means:

  1. True positives above τ: unchanged. Nothing was dropped.
  2. False positives above τ: you are seeing only 10% of them. The other 90% exist in production, they just weren't in the evaluation set.

My takeaway from the graphs is that they had a model with subsampled negatives. They evaluated performance on it, and then they evaluated performance on the bigger dataset. That seems simple enough. Where do all the funky calibration formulae come in?

Your reading of the graphs is right — and that's exactly why the formulae exist. The figure is a made-up illustration where he has both datasets in hand. The formula is for the normal case where you don't.

Why you usually can't just evaluate on the full set. Subsampling wasn't done for fun — the post opens with the reason: pipeline and training speed at 100M negatives. If you could cheaply materialize, label, and score the full negative population, you would have trained on it. The full evaluation set has the same cost problem as the full training set, plus a labeling problem: negatives must be verified safe, and at 100M messages that's not free.

So the practical situation is: subsampled eval data only, and a question about real-world precision. The formula answers it analytically instead of empirically.

Two different things are being computed, which is where the "funky" impression comes from:

  1. The precision formula — ρπ/(π(ρ−1)+1) — converts a precision measurement from subsampled to full population. This is what your graph reading does, done with algebra instead of data.
  2. Isotonic regression — maps raw model scores to calibrated probabilities on the subsampled data. This is a separate step handling the model's own miscalibration, which exists independent of subsampling.

“The precision on the full dataset will be: precision(x) = ρ (π(x) / (π(x) ( ρ-1) + 1))”

Explain this formula in simple terms.

In one sentence: it takes the precision you measured on subsampled data and asks "what would this be if I put the missing negatives back?"

Step 1 — recover the counts from the measured precision. Precision π = TP/(TP+FP). If you set TP = 1 (work per true positive), then FP_sub = (1−π)/π. At π = 0.966, that's about 0.035 false positives for every true positive.

Step 2 — restore the negatives. You kept fraction ρ of them, so the real number is 1/ρ times larger. With ρ = 0.1, FP_full = 0.35.

Step 3 — recompute precision. 1/(1 + 0.35) = 0.741. That's the 74.2%.

The formula is those three steps compressed:

precision_full = 1 / (1 + (1/ρ)·(1−π)/π)

Multiply numerator and denominator by ρπ and you get the post's version. Same thing, rearranged.

Sanity checks:

  • ρ = 1 (no subsampling) → precision_full = π. Nothing changes, correctly.
  • ρ → 0 (kept almost no negatives) → precision_full → 0. Restoring a huge hidden negative population destroys precision.
  • π = 1 (zero false positives) → precision_full = 1. If there were no false positives to multiply, scaling changes nothing.

The one thing to internalize: only the false-positive count scales. True positives are untouched because every positive was kept. All the algebra is bookkeeping around that single asymmetry.

“You can build the π using Isotonic regression on the subsampled data”

Why are we building π? Don’t we already have it? Also explain isotonic regression in simple terms in a couple of sentences.

You have π at the thresholds you measured. You need it as a function.

Measured precision comes from finite evaluation data at whatever threshold you evaluated. Three problems with using that directly:

  1. Continuity. You want precision at any τ, not just the grid points you happened to compute. Choosing an operating point means scanning thresholds.
  2. Noise. In the high-threshold region — exactly where you'd operate — few examples clear the bar, so raw empirical precision jitters. The left panel of that figure drops to single-digit sample counts near τ = 1.
  3. Monotonicity. Precision should rise with threshold. Noisy estimates violate that, giving you a curve that wiggles and a threshold choice that's an artifact of sampling.

Isotonic regression fits a smooth, monotone π(x) that fixes all three, and that function is what you plug into the ρ formula.

Isotonic regression, plainly: it fits a staircase-shaped curve to your data under one constraint — the output can never decrease as the input increases. Within that constraint it minimizes squared error, which in practice means it averages together any adjacent points that came out backwards until the sequence is non-decreasing.

Why it suits calibration specifically. It assumes only monotonicity — higher score means higher probability — and nothing about functional form. Platt scaling, the alternative, assumes a sigmoid; if the true miscalibration isn't sigmoid-shaped, that assumption costs you. Isotonic's tradeoff is that it needs more data and can overfit in sparse regions, which is a real concern here given how few positives exist at high thresholds.

Give me an example to help me understand isotonic regression. What is the input and the output?

Input: pairs of (raw model score, observed outcome). Output: a monotone function mapping raw score → calibrated probability.

Worked example. Sort your evaluation emails by raw score, bucket them, and compute the observed attack rate in each:

Raw score bucket

Emails

Attacks

Observed rate

0.1

1000

20

0.02

0.3

500

40

0.08

0.5

200

30

0.15

0.7

100

12

0.12

0.9

50

20

0.40

The 0.7 bucket is the problem — its observed rate is lower than the 0.5 bucket. That's almost certainly noise (12 attacks out of 100), not a real reversal.

What isotonic does: it walks the sequence, finds the violation, and pools the offending buckets into one by averaging them. Here it merges the 0.5 and 0.7 buckets:

(30 + 12) / (200 + 100) = 42/300 = 0.14

Both now map to 0.14. The sequence becomes 0.02, 0.08, 0.14, 0.14, 0.40 — non-decreasing. If merging had created a new violation with the neighbor below, it would pool again and repeat until the whole sequence is monotone. That procedure is called pool-adjacent-violators.

The resulting function is a step function: score 0.1 → 0.02, 0.3 → 0.08, anything in the 0.5–0.7 range → 0.14, 0.9 → 0.40. Feed it a new email's raw score of 0.6, get back 0.14 as the calibrated probability.

What changed and what didn't. The ordering of emails is untouched — isotonic is monotone, so ranking and AUC are identical. Only the numbers attached to each rank changed, from arbitrary model outputs to values that match observed frequencies.

Connecting back to the post: in his pipeline the y-axis is precision rather than attack rate, but the mechanism is the same — fit a monotone curve to noisy empirical measurements, then feed that curve into the ρ correction.

“You can build the π using Isotonic regression on the subsampled data”

Why do we use isotonic regression instead of Platt scaling?

The core difference:

  • Platt scaling fits a sigmoid: two parameters (A, B) in 1/(1+exp(A·s+B)). It assumes the miscalibration has sigmoid shape.
  • Isotonic fits any non-decreasing step function. It assumes only that higher score means higher probability.

Why isotonic fits this problem better:

  1. The distortion here isn't sigmoid-shaped. Subsampling produces a constant shift in log-odds. Composed with whatever residual miscalibration the model has (trees compress toward the extremes, nets are overconfident), the resulting correction curve has no reason to be a sigmoid. Platt would fit it approximately and leave systematic error.
  2. The extreme region is what matters, and it's where Platt is most constrained. He operates at very high precision — the post warns that a 1% error near 99% doubles false positives. A two-parameter sigmoid must trade off fit across the whole score range; it cannot bend to match the tail without distorting the middle. Isotonic fits each region independently.
  3. Data volume is not the constraint. Isotonic's usual drawback is needing more data than Platt. With 100M negatives and subsampling done for speed rather than scarcity, that objection mostly disappears.

“Online data might not match up to your batch data. Even small errors in calibrated probabilities can change the actual false positives significantly, especially if you are calibrating toward 99%. An error of just 1% will double your false positives.”

  1. What do they mean by errors here?
  2. Why would an error of 1% double false positives?
  1. An error in the calibrated probability itself — the gap between the precision you believe you're operating at and the precision you actually get. You set a threshold expecting 99%; production delivers 98%
  2. Because the quantity that generates false positives isn't precision — it's 1 − precision, its complement. And near the top of the range, the complement is tiny, so a small absolute error is a huge relative one.
    1. Believe 99% → you expect 1% of flagged items to be false positives.
    2. Actually 98% → you get 2%.
    3. 2% / 1% = 2×. Same volume flagged, twice the false positives.

“because you want to use positive examples from a different date range than your negative examples in evaluation”

Why would we want to do this?

If you want negative case freshness but positives are rare or take a while to label.

“Use online log data for calibration”

What log data are they talking about, and how would you use it?

What "online logs" means here: the production scoring record. Every email that passed through the live system, with its raw model score, the features used, the threshold in effect, and the action taken (flagged, remediated, delivered). This is emitted as a byproduct of serving — you get it for free, unlike an offline evaluation set you have to construct.

Why it beats offline data. The post's stated motivation is that (paraphrasing) online data may not match batch data. Offline sets are built from a snapshot with a chosen sampling scheme; production traffic has the real class balance, the real feature distributions, and the real drift. Calibrating on the thing you actually score eliminates the discrepancy by construction.

How you'd use it — two variants, and the second is his actual contribution:

  1. Standard. Take logged scores, attach labels, bin, fit isotonic to observed rates. Straightforward, and it fails for the reason he names: too few labeled positives, since attacks are 0.001% of traffic and labels arrive slowly from customer reports and analyst review.
  2. Volume-based. Calibrate against the count of examples flagged online rather than precision. No positive labels needed. The logic: your offline curve predicts how many messages should exceed threshold τ. Count how many actually did in production. If offline says 500/day and production flags 1,500, your calibration is off by 3× and you know it immediately — without waiting for anyone to confirm which ones were real attacks. Then use the offline precision curve to back out what precision that flagging rate corresponds to.

Why variant 2 is the clever part. It converts a slow, label-dependent signal into a fast, label-free one. Flag volume is observable in real time; precision requires adjudication that takes days. For a system doing automatic remediation, detecting a calibration drift within hours rather than a week is the difference between a contained incident and a lot of deleted business mail.

The tradeoff it accepts: flag count is a proxy. Volume can shift because your calibration broke, or because attack volume genuinely rose, or because customer mail patterns changed. The signal tells you something moved, not what — so it's a monitoring tripwire that triggers investigation, not a self-correcting mechanism.

“Calibrating against the number of flagged online examples rather than precision and then using your offline calibration to back out precision from this value is another option.”

How does this work concretely?

Explained above.

2025 Paper original

Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge

Neither offline metrics nor A/B testing are ideal. Offline metrics suffer from exposure bias and A/B testing is costly and operationally constrained.

My notes

Paper: https://arxiv.org/pdf/2508.08777

Blogpost: https://research.atspotify.com/2025/9/profile-aware-llm-as-a-judge-for-podcasts-a-better-middle-ground-between

Summary

Context

Spotify recommends podcasts to users. Evaluating podcast recommendations is typically done either through offline metrics like hit rate or recall or A/B testing.

Problem

Neither offline metrics nor A/B testing are ideal. Offline metrics suffer from exposure bias and A/B testing is costly and operationally constrained.

Solution

Profile-Aware LLM-as-a-Judge

Design Details

  • Data
    • User listening history (including podcast metadata)
  • Architecture
    • Model: GPT-4.1
    • Prompts:
      • Profile Generation: Request to generate profile following schema using provided user history (See Figure 2 in paper)
      • Evaluation: Given user profile, (1) Assess whether recommendation is good and (2) Compare two rec lists
  • Evaluation
    • Metric(s):
      • ROC-AUC, Model Selection Agreement, Recall of Strong Misalignment (Relative to ground truth from human feedback)
    • Data:
      • From 47 users included in the study, 277 pointwise human evaluations and 47 model-level comparisons (one per user). The dataset covers 227 unique recommended episodes, with an average of 5.89 episode annotations per user.
    • Results: See Table 1 and Figure 3 in paper
  • Deployment
    • N/A
  • Post-deployment
    • N/A

What I didn’t understand or am unsure about

Question

Answer

From my experience with Spotify, podcast recommendations can show up to users in different ways - they can appear on the All tab in the main feed, at the top of the Podcasts tab or in various sub-tabs. Does the paper indicate which area of recommendations they were evaluating?

It's not specified in the paper. One likely area could be a dedicated “podcasts you may like” section in the home tab.

If traditional offline metrics suffer from exposure bias, wouldn't an LLM-as-a-judge that uses synthesized user listening history to evaluate recommendations suffer from the same issue? If not, why not?

LLMs would not suffer to the same extent. When considering an item that had not previously been displayed to the user but was quite similar to items they had liked in the past, the LLM could recognize this as a good recommendation, whereas traditional offline metrics would not.

However, the profile can still under-represent latent interests the user has never been exposed to, and over-represent interests reinforced by past recommendations (a feedback-loop / filter-bubble effect, not exposure bias in the strict offline-metrics sense).

So it's not the same issue — it doesn't penalize unseen items — but it is a cousin bias: instead of "we can't evaluate what we haven't shown," it's "we can only infer what you like from what we've already shown you."

“Our two-stage profile-aware approach first constructs natural-language user profiles distilled from 90 days of listening history”

Did they indicate why they chose 90 days? If not, what are potential rationales behind arriving at this decision?

No justification of that choice is provided in the paper. Potential reasons could include:

  1. Industry convention
  2. Engineering/data-pipeline convenience (existing pipelines may already be computing 90-day features)

"This reduces input complexity and improves interpretability.”

Why does interpretability matter here?

It matters for iteration, especially given that we are collecting feedback from participants in the study about the synthesized versions of their profiles.

Any gaps in the profile identified by the participants in the study can be used to improve the LLM prompt in ways that would not be feasible if we had used the raw user listening history.

“Rather than feeding raw listening logs into a prompt, the framework synthesizes a profile that highlights topical interests, stylistic preferences, and listening behaviors. This reduces prompt complexity, increases transparency (the profile can be inspected directly), and preserves alignment with human preferences.”

They mention that the profile can be inspected directly at the benefit of this approach, but the raw listening logs can also be inspected directly. What distinction are they actually trying to draw here?

Also, in what sense does it preserve alignment with human preferences?

Direct inspection of the profile is simpler than the raw listening logs.

It preserves alignment with human preferences in the sense that the evaluation metrics display either the same or higher values for the synthesized profiles versus the raw listening logs

“Standard metrics like hit rate and recall are based on historical interaction data, which introduces exposure bias.”

How do we know exposure bias is a big enough problem that it's worth addressing?

This is a well-documented problem in the recommender systems literature outside of this paper. The paper doesn't independently validate that exposure bias is large in Spotify's own data, but I think it's a reasonable assumption.

“Traditional evaluation methods, whether quantitative or qualitative, also fall short in capturing true user satisfaction or explaining why a recommendation is relevant.”

Why is knowing why a recommendation is relevant important? How would this information be concretely used?

There are a few potential reasons:

  1. Pre-launch triage: Before an expensive A/B test, teams can read why a Judge flagged something as misaligned and decide whether that's a real problem or an artifact of a bad profile/prompt — a numeric score alone can't be sanity-checked this way.
  2. Model debugging / iteration: if Model B is "worse," rationale tells engineers whether it's failing on topic match, format mismatch, or over-indexing on collaborative-filtering signals unrelated to content — actionable for fixing the model, not just scoring it.

“Implicit feedback, such as stopping after ten minutes, can signal strong disinterest, mild curiosity, or simple distraction, making interpretation highly ambiguous.”

What does this have to do with the rest of the paragraph/section?

Traditional evaluation methods may use implicit feedback but implicit feedback can have many possible interpretations.in the context of podcasts. They state this to highlight the utility of their LLM as a judge method.

Unrelated to the blog post, does personalized recommendation typically incorporate time of day or day of week?

Yes.

  • Categorical/bucketed: e.g., "morning," "afternoon," "weekday" as discrete labels or one-hot features.
  • Cyclical numeric encoding: e.g., sine/cosine transforms of hour-of-day, so 11pm and 1am are treated as numerically close instead of far apart.

“The core evaluation task, therefore, becomes one of constructing a content hypothesis: an interpretable approximation of what the user prefers, inferred from past listening behavior.”

Is there a reason we focus on doing this content hypothesis thing in the evaluation phase as opposed to the recommendation generation phase? To the extent that this is useful during evaluation, I feel like it should be similarly useful during the actual generation of the recommendations themselves.

It seems they limit their discussion to the evaluation phase because that is the phase that they're involved in, and that's the scope of the paper.

However, one potential reason is that generating recommendations for millions of users in real time is far more latency and cost-sensitive than running an offline LLM judge on a sample of candidates after the fact.

“Models like GPT-4 show high agreement with human judgments across diverse tasks.” How is “high” agreement determined? Is it based on comparisons to intra-human agreement?

It's not made clear in the paper, but I think the answer is yes.

Some agreement metrics: typically Cohen's kappa, Spearman/Pearson correlation, or raw percent-agreement between the LLM's judgment and a human label (or majority-vote human label).

“Evaluating podcast recommendations poses unique challenges due to the nuanced, multi-dimensional nature of user satisfaction. Traditional methods typically rely on observable behavior, but in longform audio contexts, such signals are difficult to interpret”

What do they mean by "difficult to interpret"? What is the consensus on which metrics to use when evaluating recommendations of longer-form content like podcasts?

By “difficult to interpret”, I think they mean signals like partial listening are ambiguous because a single observable behavior can map to multiple, contradictory underlying intents (as previously discussed).

There is no consensus in the literature, but some common metrics include:

  • Completion rate / listen-through rate is probably the closest thing to a de facto standard for audio/video long-form content (podcasts, YouTube, streaming), used as a proxy for engagement — but as this paper argues, it's a noisy proxy at best.
  • Dwell time (raw time spent) is used as a supplementary signal, sometimes normalized by episode length.
  • Return/repeat-listen behavior (did the user come back to the same show/host) is used as a longer-horizon satisfaction proxy, since it filters out one-off curiosity clicks.
  • Explicit feedback (likes, "not interested," ratings) is used where available but tends to be sparse, which is why implicit signals dominate despite their noise.
  • A/B testing on downstream business metrics (retention, session frequency) remains the actual gold standard when feasible — offline proxy metrics are generally understood in the field to be imperfect stand-ins for this, which is consistent with this paper's own framing of A/B tests as "rigorous" but costly.

They mentioned feeding podcast metadata, including the transcript, into the LLM during the profile generation phase. Should we be concerned about the high number of input tokens and the associated costs?

The paper doesn't confirm whether profile generation uses full transcripts or truncated ones. If it's just the truncated ones, this may not be a significant concern.

Practical mitigations one would expect (not mentioned in the paper) — using only transcript snippets/summaries rather than full text, capping episodes per profile (their own ablation suggests diminishing but positive returns past 20), or pre-computing/caching profiles periodically rather than per-evaluation — are the kinds of engineering solutions a production deployment would need, but this paper is scoped to demonstrating judge validity, not production cost efficiency.

Let's say a user listened to a lot of different podcast episodes during the 90 day window considered. Does the paper indicate if all of these episodes would be included in the prompt or just a subset? If it's just a subset, how is that subset determined?

It states the profile is derived from "podcast metadata... associated with episodes and shows the user has engaged with most" — this implies some kind of engagement-based ranking/filtering, not the full 90-day catalog of everything listened to.

They have an ablation experiment at the end which could guide the choice of the subset size. There isn’t any discussion of the specific engagement metrics they used to filter though.

How did they come up with the schema included in the system prompt in Figure 2? If it's not specified in the paper, how could one feasibly come up with something like this?

This isn't specified in the paper, but it could be obtained using qualitative user research, borrowing from previous literature on similar topics or even asking the LLM itself to propose candidate dimensions. I'm guessing it was a somewhat informal process as opposed to a rigorous, separately validated design though.

In the evaluation phase, were users asked to listen to the entirety of the podcast episodes recommended? If not, should we be concerned that their initial assessment could differ from the LLM and the judges even though they would’ve felt differently about the recommendation if they’d listened to the whole thing?

They mention “a playable audio segment”, indicating that users did not listen to the entirety of the podcast episode.

I think this is a legitimate limitation, but addressing it would probably be very costly in terms of requiring significant survey participant time. It's worth thinking more about good ways to measure and mitigate if needed.

“Model A, which was primarily content-based, with less sensitivity to consumption patterns”

What is the difference between a model that's primarily content-based versus one that's more sensitive to consumption patterns?

Content-based (Model A's approach):

  • Recommends based on the attributes of the item itself — topic, description, transcript text, tags, genre, host, etc. — matched against a representation of what the user is inferred to like based on similar attributes of things they've engaged with.
  • Doesn't inherently need to know how a user consumes something (skip rate, completion, replay behavior) — it primarily needs to know what they consumed, thematically/topically.
  • This is consistent with the paper's phrasing "less sensitivity to consumption patterns" — content-based approaches are typically weaker at picking up on behavioral/format signals (e.g., "this person always abandons episodes over 45 minutes," or "prefers structured interviews over rambling banter") unless those signals are explicitly engineered in as features.

Collaborative filtering (Model B's approach):

  • Recommends based on patterns of interaction/consumption across many users — e.g., "users who listened to/finished episodes similar to what you finished also listened to X" — without necessarily needing to understand the content of the episode itself (topic, transcript) at all.
  • Inherently more sensitive to "consumption patterns" because its entire signal comes from behavioral interaction data (what was played, skipped, completed, replayed, and by whom), rather than content semantics.
  • Weaker on "content-based integration" per the paper's description, meaning Model B may be less attuned to the actual topical/thematic fit of an episode and more driven by behavioral similarity to other users.

“In model-level comparisons, the profile-based variant outperforms the history-based, underscoring the value of summarizing multifaceted user interests for reliable comparative judgments between recommendation models.”

This statement is based entirely on the MSA (W/T/L) metrics, right? If so, it seems unlikely that those differences are statistically significant, given the small sample size, right?

I think the right statistical test to use is McNemar's test. I haven't done the math, but I don't think this difference is statistically significant.

A more reasonable interpretation of Table 1, absent stats, is: the two variants perform roughly comparably, and the paper's stronger, more defensible claim is really the "comparable to raw-history despite being far more compact" framing — not a claim of Profile being reliably better.

“This tendency may be addressed through more adaptive in-context learning strategies or by model fine-tuning”

What would actually be the best way to do this? It feels like a simple prompt-based strategy where you just tell the LLM to be less decisive could work.

It's hard to say without actually experimenting with it. It would probably be helpful to be more specific, though. For example,

  • Rather than a vague "be less decisive," give the Judge concrete, structured criteria for when a tie is appropriate — e.g., "if the two lists differ on fewer than N dimensions, or the differences are minor stylistic ones rather than topical mismatches, output a tie."
  • Few-shot calibration examples: Include labeled examples in the prompt showing genuine tie cases (from the human-annotated data) alongside clear-winner cases, so the model has concrete reference points for what "close enough to tie" looks like — this is essentially what the paper means by "adaptive in-context learning."

Could also use:

  1. Confidence-based post-hoc thresholding: Have the Judge output a continuous preference score (not just A/B/Tie) and calibrate a "tie zone" threshold empirically against the human-labeled dataset — e.g., if scores within ±X of the midpoint correlate with human ties, force outputs in that band to be labeled Tie. This moves the tie-decision outside the LLM's own judgment and into a controllable, tunable post-processing step, and would directly leverage the human-annotated data they already collected.
  2. Fine-tuning / calibration on the human-labeled dataset (what the paper suggests): Directly training the model (or a lightweight calibration layer,) on human tie/non-tie patterns — the most robust fix, but the most expensive, and needs more labeled data than the 47 examples here.

“Qualitative feedback from human annotators revealed their judgments were influenced by factors beyond standard evaluation metrics”

What would be the easiest way to incorporate this feedback from the human annotators into improving the next version of the LLM as a judge?

  1. Extend the profile schema.
  2. State the judge's prompt to explicitly weigh these dimensions.
  3. Re-run the updated prompt on the existing human-annotated data to see if these improvements helped.

As described in the paper, the LLM, as a judge, just uses the user's synthesized listening history to evaluate the recommendation. Would it have been helpful to also provide the LLM with the overall most popular podcasts/episodes? I feel like human preferences typically display some level of regression to the mean.

Not sure. Maybe for cold-start users with thin profiles?

“These insights point to opportunities for enhancing profiles by incorporating long-term behavioral signals and more nuanced metadata.”

What are some specific examples of long-term behavior signals and more nuanced metadata?

Long-term behavioral signals:

  • History beyond 90 days (e.g., 1+ year) to separate enduring vs. seasonal interest
  • Show/host loyalty sustained over many months
  • Trend direction — growing/stable/declining interest in a topic

More nuanced metadata:

  • Host identity (directly named in feedback)
  • Tone/style — humorous vs. serious, scripted vs. conversational
  • Format — interview vs. monologue vs. panel, episode length
  • Listener-generated tags/reviews vs. only editorial descriptions

“Prompting LLMs with these profiles, rather than raw behavioral data, enables more accurate and interpretable alignment judgments at both episode and model levels.”

What specific metrics is the accuracy part of this statement based on? The MSA?

Yes, I think so, and as discussed previously, the statistical significance here is questionable.

How would I convince a business leader that this project was worth doing?

  1. Justify why good recommendations matter.
  2. Justify model evaluation as a concept.
  3. Highlight costs of A/B testing with online metrics.
  4. Highlight insufficiency of offline metrics.
  5. Demonstrate success of this profile-aware LLM as a judge in addressing these concerns through reduced cost and increased alignment with human preferences.

How would you improve this system if you were building it today?

1. Profile construction stage

  • Justify or tune the 90-day window empirically (currently unjustified, per our earlier discussion) — run an ablation on window length the same way they ablated episode count (5 vs. 20).
  • Add the explicitly-named missing dimensions: host identity, stylistic tone, novelty-vs-familiarity preference (all flagged by their own users in qualitative feedback).
  • Add a population/popularity signal as a variance-reduction fallback, not a universal input — specifically for thin/cold-start profiles, where their own data shows accuracy degrades (5-episode profiles: 0.51).
  • Control token/cost scaling: define an explicit, documented episode-selection and transcript-truncation policy (unstated in the paper) rather than an ambiguous "top episodes by engagement."

2. Judgment stage

  • Fix the false-positive/optimism bias (17% false positive rate, their largest reported error) via calibration against the human-labeled data — e.g., post-hoc threshold calibration (they cite Sahoo et al. [11] but don't apply it) rather than raw binary output.
  • Fix under-use of ties with structured tie criteria in the prompt or a continuous score + calibrated tie-band, rather than a vague "be less decisive" instruction (as discussed earlier).
  • Segment length disclosure/fix: since human ground truth itself is based on short audio segments (not full episodes), I'd either use longer segments or explicitly test whether short-segment judgments predict post-full-listen satisfaction — otherwise both the LLM and the human baseline share the same blind spot.

3. Evaluation/validation methodology

  • Report significance testing (McNemar's test for the paired MSA comparison, confidence intervals for ROC-AUC) — as discussed, the current headline "outperforms" claim rests on an untested 2-comparison gap out of 47.
  • Report intra-human agreement as a ceiling/baseline — without knowing how much humans agree with each other, there's no way to judge whether 0.66 MSA is "near-human" or far from it.
  • Scale the study: 47 users / 277 annotations is small for the claims being made; I'd want at least an order of magnitude more before trusting the model-level comparison specifically.

4. Scope extension (beyond what the paper attempts)

  • Test whether the same profile-based "content hypothesis" transfers to the generation phase, not just evaluation — an idea we discussed earlier that the paper doesn't explore but which their own related-work citations ([17],[18]) gesture toward.
  • Track profile drift over time and distinguish stable long-term taste from recency-driven noise (their own limitation, §4) — e.g., maintain a slower-updating "core profile" alongside a faster-updating "recent context" layer.

Level of understanding: 🤓

2025 Blogpost original

A practical framework for LLM system evaluations for multi-step processes

“Both are multi-agent systems where specialized agents handle different stages of the pipeline, each with their own failure modes. There's rarely one correct answer in either product, but there are egregious errors a human would never…

My notes

Blogpost: https://watershed.com/en-GB/blog/a-practical-framework-for-llm-system-evaluations-for-multi-step-processes

Additional context:

  1. https://watershed.com/en-GB/blog/build-evals-you-can-sustain
  2. https://watershed.com/en-GB/blog/ai-reporting-esg
  3. https://watershed.com/en-GB/solutions/product-footprints

Summary

Context

Watershed is a climate tech company that helps businesses track their greenhouse gas emissions and other environmental impacts. Two major ways it uses AI are for Product Footprints (emissions calculations for purchased products) and Report authoring (AI-accelerated environmental report generation).

Problem

“Both are multi-agent systems where specialized agents handle different stages of the pipeline, each with their own failure modes. There's rarely one correct answer in either product, but there are egregious errors a human would never make, and subtler ones that require expert judgment to catch.”

Solution

They developed a framework with two kinds of evals:

  1. Property evals: does this output satisfy known constraints?
    1. Examples: (1) do the units match (2) is the JSON well-formed
    2. Don't need input-specific ground truth so can use LLM-as-judge. Can be used in production too.
  2. Correctness evals: is this the correct answer?
    1. Examples: % difference in the emissions estimate against a reference value

Some key insights/new things I learned

  1. Distinction between property evals and correctness evals

What I didn’t understand or am unsure about

Question

Answer

“Report generation uses a RAG pipeline to generate report responses from arbitrary input documents.”

When you’re dealing with arbitrary input documents in production, what's the best way to make your pipeline robust to the different potential formats?

Ingestion/parsing: Route documents through format-specific parsers (PDF, DOCX, XLSX, scanned images via OCR) into one common intermediate representation (e.g., structured text/markdown with layout metadata preserved) so everything downstream operates on a single schema regardless of source format.

Validation/monitoring: Add property evals at the parsing boundary (did parsing succeed, is the schema well-formed) as guardrails, and continuously mine production traces for new document types the eval set doesn't yet cover — this is the direct fix for the exact staleness problem the article describes in "Is this eval truly representative of the customer problem?"

“Our evals needed to catch those while letting us iterate fast on everything else.”

What does the life cycle of product development look like with these kinds of AI products and how do evals fit into it?

{Requires additional thought and research}

Is there a structured and rigorous way to avoid overfitting as we integrate evals into the product development process?

  1. Maintain a held-out eval set that's never looked at during iteration — only checked before a release — analogous to a train/val/test split, but applied to evals themselves rather than model weights.
  2. The person iterating on the system shouldn't be the sole author of the evals they're graded against — have a second reviewer (or the domain expert, as in the article's ground-truth process) sign off on eval additions.

How is faithfulness typically evaluated?

  1. LLM-as-judge — model checks if a claim is supported by the source text.
  2. NLI/entailment models — smaller classifier labels each claim as entailed/contradicted/neutral (e.g., SummaC, AlignScore).
  3. QA-based consistency — generate questions from the claim, compare answers from source vs. output (e.g., QAGS).
  4. Citation verification — requires the generation step to cite the specific source span for each claim; a checker confirms the span actually supports the claim.

Should faithfulness always be considered a property eval or are there circumstances where it makes more sense to consider it a correctness eval?

They say later ‘"For each sentence in the report response, is this sentence supported by these cited text chunks?" An engineer does not require deep domain expertise to make this judgement.’

It’s not obvious to me why domain expertise would never be important here.

I think the blogpost relies on faithfulness being reduced to surface entailment but in more complex scenarios, that may not hold.

“LLM-as-judge property evals still require eval datasets to develop/tune the judge and prevent regression”

What do they mean by ‘develop/tune the judge’ and ‘prevent regression’?

Develop/tune the judge

  1. Write and iterate on the judge's prompt/rubric until its verdicts match expert-labeled ground truth at an acceptable agreement rate.
  2. Choose things like model choice, temperature, and output format based on which setup best replicates expert judgment on that dataset.

Prevent regression

Once the judge is tuned and deployed, that same labeled dataset becomes a fixed regression test: any time you change the judge's prompt, swap the underlying model, or modify the pipeline around it, you re-run the judge against the dataset and confirm its agreement with the known-correct labels hasn't dropped.

Does the distinction between task-level evals and component-level evals apply to property evals the same way it does to correctness evals?

I can’t think of good examples of task-level property evals so perhaps property evals only make sense as component-level.

“Task-level evals show overall system quality. When we see failures, we conduct an error analysis to identify which component is responsible.”

What would this error analysis look like in practice?

  1. Collect failing task-level cases.
  2. Trace each one through the pipeline stage by stage, inspecting intermediate outputs.
  3. Localize — the first stage whose output is wrong given its input is the responsible component.
  4. Cluster failures by root cause across the sample (rarely one component explains all of them).
  5. Prioritize root causes by frequency/severity, using the article's framing questions (cost, catchability, durability).
  6. Build an eval for the top pattern(s) so future regressions are caught automatically.

“Property evals are excellent candidates for guardrails. If the property evaluation is deterministic, then it can easily be plugged into your pipeline and broken invariants can be returned to your LLM generator to correct. If it’s an LLM-as-judge, you’ll need to first determine if the additional latency from the LLM-as-judge in your runtime is acceptable.”

Is it likely that we could use a much more lightweight LLM for property evals in production to reduce latency? It feels like the kinds of properties they are checking don't require that much model bandwidth.

Yes, likely reasonable:

  1. Task shape favors small models — these are narrow classification tasks (entailed/not, valid/invalid), not open-ended generation, so smaller models lose less accuracy here than on generation tasks.
  2. Standard practice — production guardrails commonly use small distilled models or even non-generative NLI models for exactly this reason.
  3. Still needs tuning — per the earlier "develop/tune the judge" point, you'd still need labeled data to confirm the smaller model's agreement with expert ground truth before trusting it.

For correctness evals, they say “Use online requests to inform new evals”.

What does this mean concretely?

  1. If users have the ability to override/correct, we can flag those traces and route to expert review.

“As the Report generation product matured, we discovered our customer input documents didn't match our assumptions.”

In what sense did the documents not match their assumptions?

Unclear from the blogpost but could be document format mismatches (scanned vs standard PDF), length, etc.

“As the Report generation product matured, we discovered our customer input documents didn't match our assumptions. We built correctness evals based on one type of input format, but customers wanted to upload entirely different document types.”

I understand that this is a problem but what is the takeaway here? If we expect input distributions to change, we shouldn’t invest that much in correctness evals?

Correctness evals need periodic revalidation of representativeness

How would I convince a business leader that this project was worth doing?

People use Watershed because it simplifies their emissions reporting. This simplification only occurs if the output is correct and evals are the only way to be confident in your software’s output. If we’re not measuring our performance, the customer experience will degrade and reduce their interest in using the product. {I’m not sure if there’s a good way to quantify the important though}.

Are there any relevant new developments in the space since they released their blog post that I should be aware of?

{Requires additional thought and research}

  1. Various off-the-shelf tooling we can use rather than building the Watershed-style custom pipeline from scratch. However, be careful with these.
2025 Paper original

Optimizing Query Expansions via LLM Preference Alignment

The best existing approaches for query expansion are fairly slow and don’t perform that well.

My notes

Paper Link

Summary

Context

All kinds of platforms let users search using text queries that then surface relevant results. However, users may use different language than the items they’re searching for. Reconciling this language discrepancy is often done by expanding user queries through various approaches.

Problem

The best existing approaches for query expansion are fairly slow and don’t perform that well.

Solution

Aligned Query Expansion - fine-tune a LLM using Rejection Sampling FineTuning (RSFT) and Direct Preference Optimization (DPO) to generate expansions that will enable it to surface more relevant results.

Design Details

  • Data
    • Various public datasets focused on (question, relevant source) pairs: Natural Questions, TriviaQA, WebQA, Entity Questions
  • Architecture
    • T0 encoder-decoder model
    • Rejection Sampling Fine-Tuning
    • Direct Preference Optimization
  • Evaluation
    • Metric(s): Top-N retrieval accuracy, GPU Memory Occupancy, Computational Time, Diversity (more for understanding than evaluation)
    • Data: same as data section above
    • Results: in paper
  • Deployment
    • N/A
  • Post-deployment
    • N/A

Some key insights/new things I learned

  1. Doc2Query
  2. The loss function of Rejection Sampling Fine-Tuning
  3. Diversity scores using average pairwise cosine similarity of generated expansions

What I didn’t understand or am unsure about

Question

Answer

“Beyond the computational costs, the generate-then-filter strategy also faces limitations in adaptability. Models are constrained by their inability to learn”

Why is this true? Can’t we fine-tune the relevance model in the filtering step using retrieval outcomes data?

They’re referring to the generator model, not the subsequent ranking model.

I guess the authors’ thesis is that fine-tuning the generation model is better.

“In this section, we delve into various strategies for aligning LLMs with human preferences, including reward modeling, Best-of-𝑛 (BoN) sampling, Reinforcement Learning from Human Feedback (RLHF), and Direct Preference Optimization (DPO.”

The phrase "various strategies" implies that these are mutually exclusive or at least independent, but my understanding is that reward modeling is a component of RLHF and BoN, right?

Yeah, the "various strategies" framing is imprecise. DPO is a separate strategy though.

In RLHF, “the goal is to fine-tune the LLM policy”.

What do they mean by LLM policy? Is that just the LLM itself?

Yes

“For each query 𝑞𝑖 , we generate 𝑛 query expansions … by prompting the LLM with the prompt ‘To answer this query, we need to know:’ together with the original query in front, where 𝑛 = 50 in our experiments”

Why have they structured the prompt this way?

Not really clear but I guess it's designed to bridge vocabulary mismatch between query terms and document terms, by generating expansion terms that resemble document content rather than restating the question; it might be worth experimenting with other prompts systematically.

Is the Rejection Sampling Fine-Tuning loss function basically just trying to optimize for the highest probability of getting the top ranked expansions across all queries in the training set?

Yes

Why does the DPO loss function in (9) contain the reference policy as the denominator? Does this serve the same function as in regular LLMs (to prevent the model from deviating too far from the original)?.

Yes

What does increasing or decreasing beta in the DPO loss accomplish? Why would we want to do that?

It controls how strongly the model is allowed to deviate from the reference policy in order to satisfy the preference (best-over-worst) signal.

  1. High beta reduces risk of overfitting to the ranking signal or collapsing onto narrow, repetitive outputs, but also limits how much the alignment step can actually improve retrieval performance.
  2. Low beta can produce stronger gains from alignment, but risks larger distributional shift — potentially degrading fluency/diversity, or overfitting to peculiarities of the specific best/worst pairs seen in training (especially since here each query only has one best and one worst example, not a rich preference dataset).

Is their DPO formulation the same as the original DPO paper?

Yes

“This objective ensures that the model learns to generate expansions that are more likely to be ranked higher in terms of retrieval effectiveness.”

Although the evaluation does indeed show that this works, is there a good theoretical reason that we would be confident that this strategy is worth exploring prior to any evaluation?

Since we have good ground truth labels, it stands to reason that large models can learn the relationships needed to rank well.

Why did they only use one epoch to train?

“we avoid the need for multiple sampling passes as in the filtering approach, which significantly reduces computation time”

But this sampling can be parallelized, right? So why does their approach significantly reduce computation?

  1. Batch size is capped by GPU memory, so full parallelization of all 50 samples isn't always feasible.
  2. Latency in production isn't just raw FLOPs — it includes fixed overheads (model loading, batching orchestration, inter-step data movement between generation and reranking stages) that scale with the number of steps in the pipeline, not just parallelizable compute.

The datasets they use for evaluation seem very different from the kind of queries you'd expect users to be searching on Spotify. Is that just to demonstrate the effectiveness of their approach outside of the original domain, or is there some connection between those data sets and the Spotify data that I'm missing?

This seems like more of a methods/general-IR paper rather than a Spotify-product paper. They might test on a Spotify-specific application later.

“We use this [T0] to ensure comparable results with previous work.”

T0 at this point is quite old. Isn't it worth trying some of the newer models to see how good this approach can get?

Yeah, it probably is but this paper seems more focused on demonstrating how their approach improves on existing approaches (which requires using the same model for comparability).

“fine-tune a DeBERTa V3 base as the reranker”

How do you use an encoder-only model like this one as a reranker?

“[CLS] query [SEP] candidate [SEP]” with regression head for relevance score.

Top-N retrieval accuracy is the same as recall, right?

Yes

It doesn’t seem like their methodology accounts for cases where we might have multiple relevant documents - is that a major concern?

“Major” is a strong word but it is a fair concern. Depends on the data though.

Why didn’t they use Doc2Query as one of the baselines?

Unclear.

Why didn’t they use SPLADE as one of the baselines?

SPLADE is downstream in search pipelines than where they’re putting AQE. This paper is focused on the query expansion phase, not the retrieval algorithm,

“combination of RSFT and DPO”

What does it mean concretely to combine RSFT and DPO? What’s the reference model in DPO at that point? Is it the RSFT model or the original one?

Probably RSFT followed by DPO but unclear what the reference model is - most likely it’s the post-RSFT model but not confirmed in the paper.

“In contrast, our alignment methods are more model-agnostic”

I don’t get why they say ‘model-agnostic’ here as opposed to ‘dataset-agnostic’.

I think this phrasing is imprecise.

“DPO, which explicitly optimizes for a balance between the best and worst expansions”

In what sense is this true? My understanding of their DPO formulation was that they don't want the worst expansions and their loss function is minimized when the likelihood of the worst expansions is reduced.

I think this phrasing is imprecise and I don’t think their diversity explanation is really well-supported by the evidence in the paper.

How would I convince a business leader that this project was worth doing?

You’d have to connect the relevance and latency/cost improvements to eventual plans to implement this in Spotify search. Improved search relevance and reduced latency enhances user experience, reducing churn and increasing conversion from free to paid plans. Cost reductions are $’s in Spotify’s pocket.

How would you improve this system if you were building it today?

1. Generation & ranking signal improvements

  • Stronger base model for generation: as discussed earlier, T0 (2021) is dated — swapping in a more capable instruction-tuned model for the zero-shot generation step (Section 3.1) could raise the quality ceiling of the initial 50 candidates before any alignment happens.
  • More robust ranking signal: currently rank comes from a single BM25 run per expansion (Section 3.2) — this is a noisy, binary-ish signal (as discussed earlier re: single-relevant-document limitation). Averaging across multiple retrieval runs, using a stronger/dense retriever for ranking (not just BM25), or incorporating multiple relevant documents where they exist would produce cleaner best/worst labels for RSFT/DPO to learn from.
  • Richer preference signal beyond best/worst: currently only the single best and single worst of 50 expansions are used (Eq. 5-6), discarding the other 48. A listwise or multi-pair preference objective (using more of the ranked list, not just the extremes) could extract more signal per query without needing more training queries.

2. Alignment methodology improvements

  • Iterative/online alignment: the current pipeline is a one-shot offline process — generate once, rank once, align once. An iterative loop (align, regenerate expansions with the updated model, re-rank, align again) could compound improvements, similar to how iterative DPO/RLHF rounds are used in some LLM alignment pipelines.
  • Explicit diversity regularization: since Section 5.4 shows RSFT and DPO individually reduce diversity, and only the combination recovers it, a more principled fix would be to add an explicit diversity term to the loss (e.g., an entropy bonus, or penalizing similarity to previously-generated expansions) rather than relying on the RSFT→DPO combination to incidentally produce it.
  • Address the catastrophic forgetting/staging risk — explicit checkpointing/regularization strategy for the two-stage RSFT+DPO pipeline, rather than leaving PrefP_{\text{ref}} Pref​'s identity and forgetting risk unaddressed.

3. Evaluation & deployment improvements

  • Test on production-like query distributions: validate on domain-specific query logs, not just open-domain QA benchmarks, before drawing conclusions about real-world applicability.
  • Ablate β and other hyperparameters: the paper reports only β=0.1 with no sensitivity analysis — systematically sweeping this would clarify the retrieval-accuracy/diversity tradeoff rather than relying on a single untuned value.
  • Multi-relevant-document evaluation: extend the single-document assumption (discussed earlier) to handle queries with multiple valid relevant documents, using proper recall/precision-at-k rather than binary hit-rate, for benchmarks like TriviaQA where this is plausible.

4. Query-type routing before expansion

  • Not every query benefits equally from expansion — the paper's own baseline results show zero-shot expansion sometimes underperforms the original query (e.g., Table 2, TriviaQA: zero-shot expansion trails original query on some metrics). An agent could first classify whether a query is likely to benefit from expansion at all (e.g., already well-specified vs. ambiguous/short), and skip expansion entirely for queries where it's unlikely to help — avoiding wasted computation and the occasional case where expansion hurts.

Level of understanding: 🤓

Code/Implementation

(Broad outline generated through Claude)

Platform/hardware

  • T0 3B is a mid-sized encoder-decoder model — fits comfortably on a single modern GPU (e.g., A100 40GB) for fine-tuning at batch size 16, especially with mixed precision. This isn't a massive distributed-training job; more likely single-node, possibly single-GPU or a small multi-GPU setup for the DeBERTa reranker baseline and T0 fine-tuning runs done in parallel across experiments (four datasets × three alignment methods × in/out-of-domain evaluation = a fair number of runs, but each individually modest).

Core libraries, by pipeline stage

  • Model loading/fine-tuning: Hugging Face transformers almost certainly — T0 is a Hugging Face-native model checkpoint, and DeBERTa V3 likewise. AutoModelForSeq2SeqLM for T0 (encoder-decoder), AutoModelForSequenceClassification (or a custom head) for DeBERTa V3 as reranker.
  • Training loop: Either raw PyTorch training loops, or Hugging Face Trainer/accelerate to handle the AdamW optimizer setup, batching, and mixed precision — accelerate in particular is a natural fit for keeping training code hardware-agnostic across their likely multiple GPU configs.
  • DPO/RSFT loss implementation: Since DPO isn't a built-in loss in vanilla transformers, this is likely either hand-rolled (computing log-probs from both the policy and reference model, per Equation 10, and combining via the sigmoid/log-sigmoid formula) or built using Hugging Face's trl library (Transformer Reinforcement Learning), which has a ready-made DPOTrainer class — this would be the more standard, less error-prone route rather than reimplementing the DPO loss from scratch.
  • Retrieval/ranking step: BM25 implementation likely via rank_bm25 (pure Python) or pyserini (a common IR research toolkit that wraps Lucene/Anserini, widely used in academic retrieval papers) — pyserini in particular is extremely common in exactly this kind of open-domain QA retrieval research, since it's built for corpora like Wikipedia dumps used by Natural Questions/TriviaQA.
  • Diversity metric (Section 5.4): sentence-transformers for the Sentence-BERT embeddings, then straightforward NumPy/PyTorch cosine similarity computation across pairs.

Two-stage checkpoint handoff (RSFT → DPO)

  • Concretely, this likely means: train T0 with the RSFT objective, save that checkpoint to disk (or a model registry), then load it back as both the trainable policy and the frozen reference model for the DPO stage (per our earlier discussion of what PrefP_{\text{ref}} Pref​ probably is) — trl's DPOTrainer is specifically built to take a policy model and a frozen reference model as separate arguments, which maps cleanly onto this exact staged setup.

1. Load T0 3B → generate 50 expansions/query (batch generation, top-k=50, temp=1.0)

2. Run BM25 (pyserini) retrieval per expansion → rank position of true doc

3. Select e_best / e_worst per query (numpy argmin/argmax over ranks)

4. Stage A: RSFT fine-tune T0 on (query, e_best) pairs — standard seq2seq cross-entropy

5. Stage B: DPO fine-tune (trl DPOTrainer) using RSFT checkpoint as both policy init + reference

6. Inference: greedy decode single expansion → BM25 retrieval → Top-N accuracy eval

2024 Blogpost original

Introducing Contextual Retrieval

Traditional RAG solutions remove context when encoding information, which often results in the system failing to retrieve the relevant information from the knowledge base.

My notes

Blogpost: https://www.anthropic.com/engineering/contextual-retrieval

Cookbook: https://platform.claude.com/cookbook/capabilities-contextual-embeddings-guide

Summary

Context

RAG is a very popular method to provide an LLM relevant context from a knowledge base to enhance its response quality. RAG relies on splitting documents into smaller chunks for efficient retrieval.

Problem

Traditional RAG solutions remove context when encoding information, which often results in the system failing to retrieve the relevant information from the knowledge base.

Solution

Contextual Retrieval: use the LLM to provide context for each chunk with respect to the document containing it.

Design Details

  1. Pipeline: RAG with hybrid search (semantic + BM25), as well as context labeling of chunks using the following prompt

Figure from the post

  1. Evaluation
    1. Metric: recall@20

Some key insights/new things I learned

  1. Prompt caching makes context labeling much cheaper

What I didn’t understand or am unsure about

Question

Answer

“If your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base in the prompt that you give the model, with no need for RAG or similar methods.”

Does this depend on the application, or is this broadly true?

Depends on a variety of factors including

  1. Latency sensitivity of application
  2. Consistency of knowledge base - frequently changing knowledge bases mean prompt caching is less effective
  3. Growth of knowledge base - if you expect the knowledge base to increase in size, it might be better to implement a RAG approach to preempt any problems down the line

What are the typical chunking strategies in traditional RAG?

  • Fixed-size chunking: Split text into fixed token/character counts (e.g., 200–500 tokens), often with some overlap (e.g., 10-20%) between consecutive chunks to avoid cutting off relevant context at boundaries.
  • Sentence/paragraph-based chunking: Split along natural language boundaries (sentences, paragraphs) rather than arbitrary token counts, preserving semantic coherence.
  • Recursive/hierarchical chunking: Try splitting on larger structural units first (sections, paragraphs), and recursively fall back to smaller units (sentences, words) if a chunk is still too large.
  • Semantic chunking: Use embeddings or topic-shift detection to split text where meaning changes, rather than fixed size or punctuation rules.
  • Document-structure-aware chunking: Split according to native structure (headers, markdown sections, code functions/classes, table boundaries) so each chunk maps to a logical unit.

“Developers can now cache frequently used prompts between API calls, reducing latency by > 2x and costs by up to 90%.”

Why would prompt caching reduce latency by a different factor than it reduces costs?

  • Cost savings are roughly proportional to how much of the input you avoid reprocessing — if 90% of your prompt is cached, you skip paying full price for that 90% (cached tokens are billed at a steep discount), so savings scale almost linearly with the cached fraction.
  • Latency savings depend on where the bottleneck is in generating a response. Processing the input (the "prefill" step) is only part of total response time — the model still has to generate output tokens one at a time (autoregressive decoding), which isn't sped up by caching at all. So even if you eliminate 90% of the input-processing cost, total end-to-end latency (input processing + output generation + network overhead) only improves by whatever fraction the input-processing step represented in the first place. If output generation or other overhead dominates the request, a huge reduction in prefill cost only yields a modest overall latency improvement.

Is there any good way to know if the LLM is doing a good or bad job of context labeling?

1. Intrinsic evaluation (judging the context text itself, independent of retrieval outcomes)

  • Accuracy check: Does the generated context state only facts actually present in the source document (no hallucinated names, numbers, dates)?
  • Relevance check: Does the context meaningfully situate the chunk (company, time period, section), or is it vague/generic boilerplate?
  • Consistency check: Does the same chunk produce stable, similar context across repeated runs, or does output vary widely (a sign the model is guessing rather than grounding)?

2. Extrinsic evaluation (judging effect on the retrieval/generation pipeline)

  • Ablation comparison: Run retrieval with and without contextualization, compare recall/precision at fixed K (this is what the post's own methodology does).
  • Downstream task performance: Measure whether final generated answers improve in accuracy when using contextualized chunks vs. plain chunks.

Method for scoring within either category:

  • Human review (spot-checking a sample)
  • LLM-as-judge (a separate model call scoring against a rubric)

“The resulting contextual text, usually 50-100 tokens”

Is there a reason that they haven't specified any limit on the word/token count in the prompt for the context?

Some more complex documents could require longer context descriptions to adequately capture.

Is there any way to implement prompt caching with multi-modal RAG? Let’s say the documents contained images or were scans.

Prompt caching generally works at the token level, including image tokens once an image is converted into the model's internal representation — so yes, cached prompts can include images, not just text, in APIs that support both.

“We use 1 minus recall@20 as our evaluation metric, which measures the percentage of relevant documents that fail to be retrieved within the top 20 chunks.”

How did they determine which documents were relevant in the first place?

It’s not explicitly outlined in the blog post but there are a few ways to do this:

  1. Manual/human-labeled relevance
    1. Human annotators write or review questions paired with known source passages, explicitly marking which chunk(s) contain the answer.
    2. Often used when building a QA eval set from scratch — the question is constructed from a specific passage, so relevance is relevance-by-construction (you know which chunk it came from).
  2. Automated/synthetic relevance generation
    1. An LLM is given a document/chunk and asked to generate a question whose answer is contained in that chunk — the source chunk is then automatically labeled as "relevant" for that generated question.
    2. This scales much faster than manual labeling and is commonly used when building large eval sets across many domains, though it can introduce bias (the LLM's generated questions may be easier to retrieve for than real user queries would be).

Why have they focused on recall@20 as opposed to any ranking metrics or final downstream performance?

  1. Why a recall-style metric over ranking-quality metrics (e.g., NDCG, MRR)
    1. The order of the retrieved chunks matters less in RAG than in search or recommendation since the LLM perceives them differently than a human would.
  2. Why retrieval-level metrics over downstream/end-to-end task performance
    1. Isolating retrieval failure lets you attribute cause: if retrieval fails, the generation step never had a chance regardless of model quality. Measuring downstream task accuracy (e.g., final answer correctness) conflates two different failure sources — bad retrieval and bad generation/reasoning — making it harder to isolate the effect of contextualization specifically.
    2. Downstream metrics (like answer accuracy, judged by a human or LLM-judge) are also more expensive/slower to run at scale across many configurations, since they require full generation + evaluation, whereas recall@K can be computed directly against known relevant chunks.
    3. Since their goal is testing a retrieval technique in isolation, a retrieval-level metric is the more surgical evaluation choice.

“Whereas Contextual Retrieval improves performance across all embedding models we tested, some models may benefit more than others.”

Is there any particular reason why some models would benefit more than others?

1. Model architecture/design differences

  • Context window/input length handling: Some embedding models are optimized for short inputs (e.g., single sentences) and may not encode longer, context-prepended chunks as effectively — the added contextual text could get diluted or truncated depending on how the model pools/aggregates token representations.
  • Pooling strategy: How a model converts token-level representations into a single chunk embedding (e.g., mean pooling vs. using a special [CLS]-style token) affects how much added context actually shifts the final embedding versus being averaged away.

2. Training data/objective differences

  • Pretraining task: Models trained with objectives that reward distinguishing fine-grained semantic differences (contrastive learning on diverse, hard negative examples) may be better at leveraging extra disambiguating context than models trained mainly for broad topical similarity.
  • Domain coverage in training data: A model trained on a narrower or different distribution of text (e.g., mostly short queries vs. long documents) may respond differently to having document-style context prepended, compared to a model trained on longer-form, document-style text.

“While the generic prompt we provided works well, you may be able to achieve even better results with prompts tailored to your specific domain or use case.”

Let's say I was interested in experimenting with different prompts. Is there a systematic strategy for finding the best prompt for contextualization?

1. Generating candidate prompts

  • Manual variation: Systematically vary one dimension at a time (e.g., level of detail requested, inclusion of a domain glossary, explicit instructions about what to prioritize — entities, dates, section titles) so you can isolate which change drives improvement.
  • Automated prompt optimization: Use an LLM to generate/mutate prompt variants (sometimes called prompt search or meta-prompting), scoring each against your eval set and iterating — analogous to a hyperparameter search but over prompt text.

2. Evaluating candidate prompts

  • Fix the eval set first: Use the same method discussed earlier (a held-out set of query–relevant-chunk pairs) so every prompt variant is scored against identical ground truth — this makes comparisons apples-to-apples.
  • Use the same retrieval metric consistently: Apply recall@K (or whichever metric you've chosen) across all variants rather than switching metrics mid-experiment, again for comparability.
  • Control for cost/token length: Since contextualization cost scales with output tokens generated per chunk, track average context length per prompt variant alongside accuracy — a prompt that's marginally more accurate but 3x longer may not be worth it depending on your caching/cost setup.
  • A/B or holdout testing at the downstream level: Once you've narrowed candidates using retrieval-level metrics, validate the top 1-2 prompts against downstream task performance (final answer accuracy) to confirm the retrieval gains actually translate.

They discuss using reranking models to improve results. In the context of this blog post, how does the re-ranking model differ from the ranking model? Is it just cross encoder versus bi-encoder?

Yeah, probably.

How would you improve this system if you were building it today?

  • Late-interaction retrieval models (e.g., ColBERT-style): These sit between bi-encoders and cross-encoders — token-level matching without the full cost of joint encoding — potentially reducing reliance on a separate reranking stage.
  • Hybrid dense+sparse fusion tuning: Rather than a fixed rank-fusion of BM25 + embeddings, learn query-dependent weighting (some queries benefit more from lexical matching, others from semantic).
  • Query rewriting/expansion: Use an LLM to rewrite or decompose the user's query before retrieval (especially for multi-hop or ambiguous questions), which the post doesn't address at all — it treats the query as fixed.
  • Retrieval necessity check: first decide whether retrieval is even needed (e.g., a general knowledge question vs. one requiring the document base), avoiding unnecessary retrieval calls.
  • Iterative/multi-hop retrieval: Instead of retrieving once and passing top-20 chunks to generation, an agent can retrieve, assess whether the retrieved chunks are sufficient to answer the question, and issue follow-up retrieval queries if not — useful when the first retrieval surfaces partial or tangential information.
  • Tool-augmented retrieval: An agent could choose between multiple retrieval sources/tools (e.g., a structured database query vs. the vector/BM25 knowledge base) based on the nature of the question, rather than always routing through the same pipeline.
  • Self-critique/re-retrieval loop: After generating a draft answer, an agent can check whether the answer is actually supported by the retrieved chunks, and trigger additional retrieval if it detects unsupported claims — this connects to the "always run evals" note in the post, but as a runtime check rather than an offline eval.
  • Citation grounding checks: Verify each claim in the final answer traces back to a specific retrieved chunk, flagging or re-retrieving for any that don't.
2026 Blogpost original

How we Built Image Understanding for Legal Documents

Image processing is expensive but visual content is often needed to answer the user’s question. How do you balance these two priorities?

My notes

Summary

Context

Harvey has a legal AI agent that lawyers can use to answer questions about documents. Documents can consist of both text and images.

Problem

Image processing is expensive but visual content is often needed to answer the user’s question. How do you balance these two priorities?

Solution

On-Demand Analysis: the agent recognizes when a user asks a question that involves visual content and invokes the image analysis tool.

Some key insights/new things I learned

  1. Processing a single image is roughly 50x more expensive than generating a text response

What I didn’t understand or am unsure about

Question

Answer

On a general note, are image components and text components in a PDF file clearly distinguishable from the file content on its own? Let's say I want to extract all images on a specific page in a PDF - can standard PDF parsers do that deterministically?

Text and image content in a PDF are structurally distinguishable: text comes from text-showing operators referencing fonts, while raster images are explicit Image XObjects (or inline images) with their own dictionaries and compression filters. This means a standard parser can deterministically enumerate embedded raster images on a given page by walking the object structure — no guessing required.

However, this doesn't capture everything a person would call "an image." Vector graphics like charts and diagrams exported from tools like Excel or PowerPoint are drawn as paths/curves rather than stored as discrete image objects, so they won't show up if you only extract Image XObjects — you'd need to render the page to see them. Scanned pages may be a single full-page image with no separate text layer, and resources can be nested or shared across pages, complicating "what's on page N" queries. So: finding embedded raster images is deterministic; reliably capturing all visual content (including vector charts) generally requires rendering the page rather than object-level extraction.

“Traditional document processing pipelines only extract text and have no way to extract or pass on visual information. That works for text-based contracts and memos, but falls apart when documents contain … Scanned documents with handwritten annotations”

I feel like an intelligent OCR system could handle this kind of document. Is that not the case?

  1. Handwriting recognition is still meaningfully less accurate than printed-text OCR.
  2. OCR text extraction typically doesn't preserve spatial relationships between text (for example, arrows).

They also mention traditional document processing pipelines struggling with “tables with complex layouts that don't survive text extraction”. Is there no easy way to extract these deterministically? I can understand why OCR wouldn't work. A table in a standard structured PDF seems like it’d just be a block of code on the backend.

Actually, structural relationships are never explicitly encoded. Only the visual positions are. What you have on the backend is just a content stream full of positioned text-showing operators — each one says "draw this string at this (x, y) coordinate" — plus possibly some vector lines for borders. There's no native <table>, <row>, or <cell> object in standard PDF.

So when a text extractor pulls the content out, it typically does one of two things:

  1. Reading order by position/stream order — it walks through the text-showing operators roughly left-to-right, top-to-bottom. For a simple single-column table this often works fine, since the natural stream order matches the natural reading order.
  2. Heuristic clustering — fancier extractors try to infer rows/columns by clustering text by x/y coordinates and looking for the gridlines.

This wouldn't work well for merged cells or multi-column tables or no grid lines.

“In addition, the model has no way of knowing what was dropped or left out during document processing”

Why is this the case? Is there no clean way of ensuring the model knows what was left out?

Not fully sure but my guess is that it can get pretty complicated.

“When a user asks a question in Harvey that involves visual content, the agent recognizes this and invokes the image analysis tool”

How exactly does the agent know a question involves visual content?

The model has a list of tools alongside their description.The image analysis tool. One possible component of the description could be something like “use this for questions about charts/diagrams/visual layout that text alone can't answer”. So if a user asks a question about any of those elements, the model might realize that it needs to use image analysis.

Let's say a user asks a question about some content or piece of information that's only present in an image and is not discussed at all in the extracted text. Could the existing pipeline, as described in the blog post, be able to help the user answer their question?

If the information is genuinely isolated in the image with no textual trace anywhere in the document, the architecture as described would likely struggle to locate the right page.

What would the “Smart page finding” pipeline look like in practice? Isn’t there a meaningful risk of false negatives?

  1. Hybrid search: embedding-based semantic search and lexical/keyword search (e.g., BM25 or similar).
  2. Yes, there seems to be a meaningful risk of false negatives. This could be mitigated through fallbacks such as expanding the search, lowering a similarity threshold, or full-page brute-force scanning for when candidate-page selection comes up empty or low-confidence.

What is an “oversized image”? Isn’t image resolution in PDFs capped?

No, it's not capped.

“intelligent downscaling for oversized images”

Doesn’t that risk some content becoming too small to parse?

The risk is there but I guess they’ve determined the tradeoff is worth it. Maybe you can mitigate it by content-aware scaling.

“If the agent can find the answer in the text, we don't spend more token on vision”

What does it mean to “find the answer”? What mechanism does the agent have to know if it’s found the answer?

No such explicit mechanism is described in the blog. They could be using some sort of separate verification step, confidence threshold, or completeness detector but it's not clear from the text itself. We could also surface sourcing to the user - this way, a lawyer could decide if the model was supposed to use an image or not and evaluate whether the answer is good in light of whether it pulled the appropriate sources.

“90% of those images are not actually necessary for answering the query”

How do you determine this?

Consider the previous existing pipeline. Compare the number of images that are retrieved when answering the question versus the total number of images processed.

“The tool reports what it can and can't determine, with confidence levels.”

What do they mean by confidence levels here? Are they using any kind of calibration on these confidence levels? Should they be?

It's not exactly clear from the blog post.

They should probably be using some sort of calibration. This can be evaluated against "a curated dataset of real legal queries,". It would be straightforward (in principle) to bucket the model's stated confidence levels and check, for each bucket, what fraction of those answers were actually correct against ground truth — and if the buckets don't line up (e.g., "exact" readings are only right 70% of the time), recalibrate the prompting or add a verification step.

“Foundation models are surprisingly good with just text.”

This feels less like an insight about the foundation model capabilities and more about the nature of the questions relative to the documents, right?

“It shouldn’t over trigger and affect other tool recall metrics.”

Does over triggering affect other tool recall metrics because only one tool can be triggered?

Sort of but not exactly - some agentic frameworks have a limited number of reasoning steps/tool calls in which case an incorrect image analysis tool call would consume the budget. Also, if the agent calls the image tool and gets back something that looks like an answer, it may terminate its reasoning loop there rather than continuing to check whether another tool was actually more appropriate — i.e., the agent "settles" once it gets a plausible-looking result, regardless of whether more tools were technically available.

They mention that tool descriptions are a critical component of optimizing the pipeline. What would the process of optimizing the tool descriptions actually look like?

A typical iterative loop would likely look like:

  1. Build a labeled eval set — queries labeled as needing vision or not, including tricky negatives (e.g., a query that mentions a chart but is answerable from a caption).
  2. Write an initial description, run the agent on the eval set, and log trigger decisions.
  3. Compute precision/recall on triggering — of queries that should've triggered the tool, how many did (recall); of queries that did trigger it, how many should have (precision).
  4. Error-analyze misses: false positives often point to overly broad wording (e.g., "questions about data" catching text-answerable cases); false negatives often point to missing scope (e.g., only mentioning "charts," missing diagrams or signature pages).
  5. Revise wording accordingly — tightening, adding examples, excluding known false-positive patterns.
  6. Re-test and iterate until precision/recall hit an acceptable tradeoff.
  7. Possibly A/B test in production, since an offline set may not capture real query diversity.

“In our evaluations, the system reliably extracts specific numeric values from charts, identifies visual elements in complex layouts, and provides structured answers that lawyers can cite.”

What would the evaluation metrics for each of these look like?

1. "Reliably extracts specific numeric values from charts"
This is a factual-accuracy task with a checkable ground truth, so the natural metrics:

  • Exact-match accuracy: extracted value == ground truth value (for labeled/exact readings).
  • Tolerance-based accuracy: extracted value within some % or absolute error of ground truth (for interpolated/approximate readings, given the post's "approximate vs. exact" distinction).
  • Numeric error metrics (MAE/RMSE) across all chart-value extraction tasks, separated by exact vs. approximate cases.
  • Possibly precision/recall on whether it correctly identifies which value the question is asking about (e.g., right bar/series) before grading the value itself.

2. "Identifies visual elements in complex layouts"
This sounds more like a detection/localization task:

  • Precision/recall (or F1) on whether the correct element was identified as present (e.g., did it correctly find "the signature block" or "the floor plan legend").
  • Localization accuracy, if relevant — did it correctly attribute the element to the right page/region.
  • Possibly classification accuracy if "identifies" means categorizing element type (chart vs. diagram vs. table vs. signature block).

3. "Provides structured answers that lawyers can cite"
This is the vaguest of the three, and "citable" is a qualitative/usability property rather than a single accuracy number. Plausible metrics:

  • Human/expert rating (e.g., lawyers or annotators scoring answers on a Likert scale for clarity, structure, citability).
  • Format compliance — does the output reliably include source attribution (page number, document location), since "citable" implies traceability.
  • Faithfulness/groundedness scoring — does the structured answer accurately reflect what's actually in the source, distinct from raw numeric accuracy (more relevant to qualitative descriptions than numbers).

How would I convince a business leader that this project was worth doing?

The value proposition of Harvey is that it makes lawyers better at their job. In this blog post, they focus specifically on how image understanding helps in answering user questions.

If we're talking about the image understanding system in general, we could evaluate its presence versus its absence on metrics related to the quality of answers generated.

If we're talking about the on-demand system in comparison to the original system, the metrics of interest would be focused on reduction in costs while maintaining other quality metrics.

How would you improve this system if you were building it today?

1. Fix text-first false negatives — add a confidence check on text-only answers (not just vision ones), and force vision when extraction artifacts ("see chart above," broken tables, placeholders) suggest missing visual content, rather than trusting generation success alone.

2. Fix candidate-page false negatives — when text search returns empty/low-confidence but the question implies a visual answer, fall back to broader page scanning rather than relying solely on text-match retrieval.

3. Calibrate confidence levels — LLM self-reported confidence is typically poorly calibrated; build a held-out set, measure actual accuracy per stated confidence bucket, and correct accordingly. Highest-leverage given the "lawyers can cite" stakes.

4. Disaggregate evaluation metrics — track tool-trigger recall, page-selection recall, element-detection F1, and extraction accuracy (exact vs. approximate) separately, instead of one bundled "reliably extracts" claim.

5. Smarter downscaling — detect dense small-text regions and preserve resolution there (or tile separately) instead of one global downscale factor.

6. Default source attribution — cite page/region for every answer, giving lawyers a cheap sanity-check beyond confidence labels.

7. Production monitoring, not just offline eval — track live trigger rates and user-correction signals to catch description drift the eval set missed.

8. Stress-test the no-textual-anchor case — build an eval subset of genuinely isolated visual content (no caption/reference) to measure outright failure rate, since this seems like the most structurally fragile gap.

2023 Blogpost original

Reddit’s LLM text model for Ads Safety

Reddit has a ton of posts; balancing accuracy, latency and cost in risky content classification is difficult. Reddit already had logistic regression and gradient boosted models in production but these were imperfect.

My notes

Summary

Context

People can post in various subreddits. Some of these posts might be sexually explicit or violent. Advertisers may not want their ads to show up next to this kind of NSFW content.

Problem

Reddit has a ton of posts; balancing accuracy, latency and cost in risky content classification is difficult. Reddit already had logistic regression and gradient boosted models in production but these were imperfect.

Solution

RoBERTa (a BERT spinoff) fine-tuned on labeled data.

Some key insights/new things I learned

  1. Efficient sampling through filtering model.
  2. CPU optimization frameworks

Design Details

  • Data
    • Posts labeled as X (Explicit) or V (Violent) by human labelers according to industry standards and Reddit’s policy guidelines.
    • 250k annotated samples. To counteract class imbalance from the rarity of positive samples, a filtering model trained on an open source dataset was used as a sampling filter.
  • Architecture
    • RoBERTa generates embeddings of posts that are then passed into a simple classifier (like a single-layer neural network) which predicts the label for the text.
    • Post text truncated to 4096 characters, subword tokenizer (WordPiece) and split into pieces of 256 tokens
  • Training
    • Initialize with weights shared by a sibling team.
    • Training (~80%), Validation(~10%), and Test (~10%). The training set is what the model is trained on. Each epoch, the trained model is evaluated against the validation set which is not seen during training. The set of weights that perform best against the validation set is the model we select for offline evaluation against the test set
  • Evaluation
    • Offline
      • Metric(s): accuracy
      • Data: test set
      • Results: Not shared
    • Online:
      • Metric(s): Not shared
      • Data: A/B test
      • Results: Not shared
  • Deployment
    • The model was served on a CPU (since GPUs weren’t available for online inference at the time of writing).
    • They increased the number of CPU cores in the deployment request and the parallelism (number of threads). This resulted in a further reduction in model latency due to allowing for parallel processing to take place during heavy computation operations (self-attention).
    • Used various optimization libraries (TorchScript, BetterTransformer and ONNX) and found that ONNX in particular decreased latency by a lot.
  • Post-deployment
    • Not shared

What I didn’t understand or am unsure about

Question

Answer

They say they truncate posts to 4096 characters and split them into pieces. Later, they talk about using token lengths of 512 or 256 as the piece size for inference. But they also say they use WordPiece tokenization which is not a character-level tokenizer. What’s going on here?

The 4096 characters are tokenized using WordPiece into < 4096 subword tokens, which are then split into pieces of size 256 or 512 for inference.

Once they split a post into pieces, do they just average/sum the embeddings before the final classification?

Unclear but a few approaches include:

  1. Max pooling over chunk predictions (useful for detecting if any one piece was harmful)
  2. Average pooling over chunk probabilities (probably not ideal since harmful chunks would be diluted by less harmful ones)
  3. Hierarchical approach — run a second model over the chunk-level embeddings to produce a post-level prediction

How much of an impact does initializing with somewhat informative weights (like they did) have relative to initializing with random weights?

Need less labeled data to converge to a good model so reduced cost.

Some risks:

  1. Inherits any biases present in the other model
  2. Anchoring to the wrong distribution: if the other team’s model had a completely different objectives

“When we batch embedding vectors together, they all need to be the same length for the model to properly perform the matrix computations, therefore a large batch size requires padding all embedding vectors to be the same length as the longest embedding vector in the batch. When batch size is large, then more embedding vectors will be padded which, on average, increases n. When batch size is small, n on average will be smaller due to less need for padding and this reduces the driving factor of the computational complexity.”

  1. Does padding affect prediction accuracy for the padded vectors?
  2. I understand why n on average will be smaller but it’s not clear why that is more important than reducing the number of different batches? Does the latter offer no major benefit when using CPUs?
  1. No, padding shouldn't affect prediction accuracy, assuming it's implemented correctly. Padding tokens are masked during the attention computation so the model ignores them — they don't influence the embeddings or the final classification.
  2. On CPU, you have far fewer cores, so the parallelism benefit of batching is much weaker. What dominates instead is the raw size of the matrix operations, which is driven by n (the padded sequence length). So on CPU, the cost of inflating n through padding outweighs the benefit of fewer batches, reversing the GPU calculus.

If you can just use more cores and threads on CPU, why bother with GPUs?

  • Sheer number of cores — a high-end CPU has maybe 64 cores; a high-end GPU has thousands to tens of thousands of simpler cores. For the massively parallel matrix operations that dominate transformer inference, this difference is enormous.
  • Memory bandwidth — GPUs have much higher memory bandwidth, which is often the actual bottleneck during inference rather than compute.
  • Specialized hardware — modern GPUs (and TPUs) have dedicated tensor/matrix multiply units (e.g. NVIDIA's Tensor Cores) designed specifically for the operations transformers use. CPUs are general-purpose and lack this.
  • Cost — scaling CPU cores to match GPU throughput would require many more machines, likely costing more in aggregate.

“Both TorchScript and ONNX frameworks work better without batching the inputs (i.e. running all inputs sequentially). This is likely due to reduced tensor size during computation since padding would not be required.”

Why is this true? In particular, why does this hold for these two frameworks but not for PyTorch?

TorchScript and ONNX both compile/optimize the model graph ahead of time into fixed computational graphs. Base PyTorch uses a dynamic computation graph.

This might explain the difference but I don’t fully understand why.

“Regarding model classification improvements, we have millions of Reddit posts being created daily that require us to keep the model up-to-date as to avoid model drift.”

How would you actually do this?

Detecting drift:

  1. Monitor the distribution of model output scores over time. If the average confidence or label distribution shifts, that's a signal.
  2. Track disagreement between the model and human reviewers on a sample of posts.

Retraining:

  1. The simplest approach is periodic retraining (e.g. monthly) on a sliding window of recent data, so the model sees current language and content. To avoid catastrophic forgetting, we can
    1. Mix old and new data
    2. Use a low learning rate
    3. Elastic Weight Consolidation
  2. If we have some drift detection in place, we can use that for off-cycle training.

Labeling new data

  1. Active learning — prioritize labeling examples the model is most uncertain about, which is more efficient than random sampling

What would be an acceptable latency level for new posts?

It depends on how the classifications are used downstream. We could prevent ads from appearing next to posts that haven’t been classified yet which would let us relax our latency requirements a bit.

We also have the other simpler models that can be used as a temporary classification.

What happens if the model fails to generate a prediction or crashes for whatever reason? Is it reasonable to have a backup model or are there better strategies??

Apparently, Kubernetes already handles instance failures. Otherwise, a backup model is also a reasonable strategy.

What would be a reasonable backfilling strategy (review old posts that didn’t go through the new model)?

They have GPUs they can use for batch jobs which would make this easier. For prioritization, maybe weight based on traffic, recency and risk-heuristics.

What about combining predictions from the gradient boosted model and RoBERTa?

This could be useful. Ways to combine them include:

  1. Simple averaging of output probabilities
  2. Weighted average
  3. Stacking: train a third model on the outputs of both as features
  4. AND/OR logic — e.g. only flag if both models agree (high precision) or flag if either flags (high recall), depending on whether false positives or false negatives are more costly

How would I convince a business leader that this project was worth doing?

If the new model reduced false positives all else equal, it should be relatively easy to translate that into increased advertising revenue.

If it reduces false negatives, we may need to think about how to communicate the value of the advertiser relationships (why should Reddit care if ads show up less than before alongside objectionable content).

How would you improve this system if you were building it today?

Model architecture improvements

  1. Instead of RoBERTa, use DeBERTa or an instruction-tuned LLM (like quantized Llama),
  2. Multi-task learning across multiple safety dimensions.
  3. Using a large LLM as a labeling assistant. Use the LLM's soft output probabilities rather than just its hard labels to train a smaller production model.

Multimodal coverage

  1. Use a vision-language model like Qwen2.5 or Gemma 3.

Level of understanding: 🙂

2025 Blogpost original

Elicit Reports

How do you evaluate this capability?

My notes

https://elicit.com/blog/introducing-elicit-reports/

Summary

Context

Elicit aims to automate time-consuming research tasks. One capability it offers is providing detailed reports about research questions based on systematic reviews of relevant literature.

Problem

How do you evaluate this capability?

Solution

Developed criteria and recruited researchers to evaluate Elicit and competitors against that criteria.

Some key insights/new things I learned

  1. Wilcoxon signed-ranks test

What I didn’t understand or am unsure about

Question

Answer

“Elicit will suggest ways to clarify the question or explore additional angles.”

How does it do this? Is the LLM just prompted in the background to rephrase the question as needed? Could few-shot examples be useful? Would fine-tuning be helpful or is that overkill?

Probably prompt-based with few-shot examples. Fine-tuning is probably overkill.

How do Elicit Reports (and other similar Deep Research tools) actually work? What are all the ways they differ from a regular chatbot LLM?

Elicit

  1. Elicit runs discrete steps: paper search, screening, data extraction, then report generation
  2. Users can inspect and edit each intermediate step
  3. It uses the same engine as their existing Systematic Review workflow

General

  1. Agentic loop — rather than one inference call, the model runs multiple calls, often deciding what to do next based on prior outputs
  2. Tool use — the model can call external tools (web search, database queries, PDF readers) mid-task
  3. Long context / memory management — they need strategies to handle more text than fits in a context window, e.g. chunking, summarization, or retrieval
  4. Multi-step planning — the model may decompose a task into subtasks and execute them sequentially or in parallel
  5. Retrieval over real sources — unlike a chatbot drawing on training data, these tools fetch live documents and ground claims in them
  6. Structured intermediate outputs — e.g. extraction tables, screening decisions — rather than just free text

“Elicit cites every claim with exact quotes from the original source.”

How does this feature work on the backend? How can we guarantee no hallucinations or misquotes here?

How it likely works:

  1. The source PDFs/text are chunked and stored (likely with embeddings for retrieval)
  2. When generating a claim, the model is prompted to pull the supporting quote directly from the retrieved source text that's in its context window
  3. The quote is then displayed alongside a link back to the specific paper

Ways to reduce errors:

  1. Post-hoc verification: a separate pass that checks whether the quoted string actually appears in the source document (string matching, not LLM judgment)
  2. Keeping source text in-context rather than relying on the model's parametric memory

“You can also extend a report by chatting with it. Ask questions about specific claims and papers, or have it summarize certain report sections. You can even ask it about the underlying information, even if it wasn't covered directly in the report.”

Since the report used lots of different sources, what’s the best way to have the LLM retain the ability to access all of that knowledge without overwhelming the context window and increasing cost/latency?

The core problem is that all source papers can't fit in the context window, so the system needs to selectively retrieve relevant content at query time. The standard approach is embedding-based RAG: chunk documents, embed and store them, then at query time retrieve only the chunks most semantically similar to the user's question.

The challenge is that the user’s question may be sparse/lack sufficient context to be useful as an embeddable chunk for retrieval. This can be mitigated in a few ways:

  1. Query expansion/enrichment before embedding using relevant context
  2. HyDE (Hypothetical Document Embedding) — instead of embedding the question at all, you prompt the LLM to generate a hypothetical answer, then embed that. The intuition is that a hypothetical answer lives in the same embedding space as actual source chunks, so it retrieves better matches than a question would.
  3. In a multi-turn chat about a report, each follow-up question may only make sense in light of prior turns. So the retrieval query should probably be constructed from the full conversation context, not just the latest message — though this adds cost.

What is the difference between “Was it accurate?” and “Were its claims supported by the literature?”?

I think “literature” here might mean from the cited literature. A claim could be accurate but unsupported (the model stated something true but didn't cite a source for it), or supported but inaccurate (the cited source doesn't actually say what the model claims).

Is there any indication whether the competitor Deep Research were told during the testing process that the use-case is for scientific research and not generic deep-research?

I think a prompt-based reframing has a reasonable likelihood of improving performance and so it’d be worth reporting to avoid bias.

It’s unclear from the blogpost.

How did they decide on recruiting 17 researchers, as opposed to more or less?

In theory, you could do some sort of statistical power calculation but that requires a lot of assumptions and it’s not clear from the blogpost whether that was considered. Perhaps they worked backwards from whatever budget they were assigned for this.

In the report itself, they have a “Full text retrieved column”.

Figure from the post

In cases where the full text wasn’t retrieved, why is that? Is it because Elicit didn’t have access to it or because Elicit thought it wasn’t useful?

Probably the former but unsure.

We can’t preclude the possibility of bias in the evaluations since Elicit is paying the researchers. Is there any good way to calibrate and adjust for it though?

Not sure.

“we obtained modest sample sizes of around 25-30 evaluations for each competitor”

Why is this different from the 17 number earlier?

Some researchers may have evaluated tools on more than 1 research question. Unclear whether this lack of independence in observations is accounted for in statistical significance computations.

“All p-values were calculated on direct paired comparisons of Elicit vs. a competitor tool using the same research questions”

Are these p-values adjusted to account for the multiple testing problem? Should they be?

I think so but it depends on how you frame the inference goal. Worth thinking more about.

What is the difference between “usefulness” and “how well it answers the research question”?

Intuitively, "how well it answers the research question" seems more narrow and objective — did the report actually address what was asked? "Usefulness" could be broader, capturing things like whether the output saves time, is well-organized, or surfaces unexpected insights even beyond the direct question

“Evaluators noted that Elicit relies on trustworthy academic sources, unlike competitors.“

Am I right to think this doesn’t seem like a long-term competitive advantage? If I prompted a competitor's deep research tool to use only trustworthy academic sources, I feel like it would generally successfully avoid news articles, interviews, etc. Or to the extent that it doesn’t right now, this seems like something that would be fixed soon enough.

Yeah, I can’t think of a good counterargument.

“For example, users can add their own papers, override screening decisions, and add or remove extraction questions, and then regenerate their report.”

In these circumstances, would the report be generated from scratch?

If a user just adds one paper or overrides one screening decision, rerunning the entire pipeline wastes compute on steps whose inputs haven't changed. A smarter implementation would only rerun downstream steps from the point of intervention (e.g. overriding a screening decision only reruns extraction and report generation, not the search). Whether Elicit does this isn't stated.

“We paid the evaluators, which may have biased their results. However, evaluators (a) did not hesitate to rate competitors more highly at times and (b) justified all their ratings, which leads us to believe that this is not driving results.”

These arguments don’t really seem compelling. X + b < Y if X is much lower than Y. Justifications can easily be rationalizations of bias. Am I missing something?

Yeah, I can’t really think of anything that would make this more persuasive.

Additional References

  1. https://elicit.com/blog/systematic-review/
  2. https://elicit.com/review/9a02f3e6-95ce-4a7c-8485-6c1eaee14c99?ref=blog.elicit.com
2026 Blogpost original

Financial Benchmarks

My notes

Summary

Context

LLMs are used in Ramp's products in a variety of ways. Ramp wants to make sure it can easily compare the performance of different approaches/models across these use-cases. To facilitate this,

Task

Metrics

Contextual Invoice OCR

Customers upload invoices and Ramp extracts relevant info. Customers can edit afterwards to fix any errors.

Perfect Extraction Rate (% of invoices with 0 edits required)

Financial Statement OCR

Customers upload financial statements and Ramp extracts relevant info. Less standardized than invoices and also includes calculation of metrics not directly listed.

Match Rate (% of extracted values within 1% of ground truth)

Policy Agent

Ramp has built a policy agent to determine whether expenses are in policy. This policy agent can return Approve, Unsure or Reject.

Disagreement Rate (% of transactions for which the model disagrees with the human decision)

Unsure Rate (% of transactions for which the model outputs unsure)

Accounting Autocoding

Assign expense to the appropriate category — e.g. Marketing Spend, Travel, Research and Development, Mileage / Fuel.

Accuracy@1

Partner Restrictions Compliance

Ramp partners with various financial institutions to enable its core offerings; these institutions have compliance requirements which will affect which products each customer is eligible to be onboarded onto. They built a framework which determines for each customer and each partner which restrictions, if any, apply.

Compliance Adherence Score:

  1. Correctly identifying restricted businesses (avoiding false negatives).
  2. Covering the full breadth of restriction categories.
  3. Avoiding false positives.

Fund Smart Routing Agent

Ramp uses an LLM agent to choose the most appropriate fund to route a transaction to.

Override Rate (% of cases routed incorrectly)

What I didn’t understand or am unsure about

Question

Answer

“We built out contextual OCR to learn patterns from a business' past submitted bills and use them to improve extraction quality”

Does contextual OCR generally consist of a lot of heuristics like the example they give or are there other approaches that could require less manual effort to set up?

One automated way to learn from previous bills would be to

  1. Construct embeddings for each bill processed in the past (these could be OCR text embeddings, VLM embeddings like ColPali)
  2. Do RAG when passing the current bill through the LLM (dynamic few shot prompting)

“we take a look at what percentage of invoices each model is able to extract perfectly”

Why do they look at the rate of perfect extraction as opposed to the average % of fields correctly extracted?

The ultimate value in accounting automation is Straight-Through Processing. Any human intervention, regardless of size, is a major cost.

Using the average % metric also elides differences between a model that is 100% correct 90% of the time (pretty useful) vs a model that’s 90% correct 100 % of the time (not that useful).

“Match Rate (% of extracted values within 1% of ground truth)”

The 1% part feels a little weird. A model that mistakes the first digit would be much more heavily penalized than a model that misses the last digit, even though it’s unlikely that they’re deterministically faulty in different ways. Is it just for simplicity, because they’re more focused on derived/calculated items or something else?

Ramp is testing LLMs rather than just OCR engines; there is a reasoning component to calculating the final metrics and so a 1% margin of error is useful as a metric.

I still feel like getting the first digit wrong vs getting the last digit wrong seems kind of random, and not innately linked to the model's capabilities. But I guess that’s just one relatively minor issue.

“Given these limitations, this problem is tricky to evaluate given that we lack clear ground truth.”

Why do they say they lack clear ground truth? I’m sure human reviewers are indeed imperfect but that’s true for any situation where you have human-labeled data - is there any reason why it’s more true here?

Not fully clear but I guess this task is more subjective than most.

For the Policy Agent, they cite a tradeoff between Unsure Rate and Disagreement Rate. If I wanted to calibrate the models to make them focus more on one goal vs the other, what would be the easiest way to do that?

  1. Prompt-based: telling the model to lean one way or another
  2. Confidence Thresholding: ask the model to output confidence scores for each possibility and then set thresholds however you want.
  3. Few-shot examples: your choice of examples can determine whether what a model is more likely to output

“Changing patterns vs one-off exceptions - new categories are often introduced as companies grow”

What would be the best way to make your autocoding pipeline robust to the addition of new categories?

Instead of telling the model "Choose from Category A, B, or C," you store all category names and detailed descriptions in a vector database. When a transaction comes in, you retrieve the most semantically relevant categories from the database and feed only those into the LLM prompt. When a company adds a new category, you simply insert it into the vector database. The model "discovers" it during the next retrieval without any code changes.

“Success in this task is defined by … accurate semantic understanding + reasoning when historical patterns are wrong or absent in a cold start”

How can you know when historical patterns are wrong? Do you just tell the LLM to rely on historical patterns unless they’re very confident those patterns are wrong?

Yeah, I think some sort of confidence score + thresholding would be the move here.

What does category coverage mean in the Compliance Adherence Score? How is it measured?

Macro-Average: take average of recall for each category and take the average of that average (this ensures each category is weighted equally)

“Identifying restricted businesses is the most vital component because a miss … would be a compliance failure.”

Let’s say I was working on this project and I wanted to calibrate the model appropriately to capture the higher priority we want to assign to false negatives vs false positives. How precisely would I do that in the model building and evaluation phases respectively?

  1. Prompt-based
  2. Confidence score and set thresholds based on desired FPR vs FNR
  3. Fine-tuning with weighted cross-entropy
2023 Blogpost original

An Abnormal Approach to Machine Learning: Feature Systems and Language Models

My notes

What I didn’t understand or am unsure about

Question

Answer

“Per-customer behavioral baselines let Abnormal flag anomalies without needing to predict attacker tactics in advance.”

How exactly do they do this?

Behavioral anomalousness can be considered in two ways:

  1. As features input into a larger model
    1. For example, cosine distance of current email to sender's typical content
  2. An independent anomaly detection model like Isolation Forest or Autoencoder (uncalibrated so difficult to use on their own)

“Abnormal uses Google's BERT model to detect new attack classes and identify attacker intent across email bodies and headers.”

How does BERT map to the outputs specified in that sentence?

Mapping each output to a plausible mechanism:

1. New attack classes. BERT is not a novelty detector; nothing in a transformer flags "I haven't seen this before." Two mechanisms actually deliver this:

  • Generalization — semantically similar text maps to nearby embeddings, so a reworded variant of a known attack scores like the original even though no signature matches. This is the polymorphic case, and it's real.
  • Campaign clustering — grouping variants by embedding proximity surfaces a coordinated campaign without a prior label for it.

Neither detects a genuinely novel class. Catching an attack type never seen in training is where the per-customer behavioral baseline does the work, not BERT.

2. Attacker intent. This is straightforward supervised classification: a head over the encoder trained on labeled examples, outputting categories like payment redirection, credential harvesting, or gift-card fraud. "Intent" is a label taxonomy, not an emergent property. BERT's contribution is that intent is expressed semantically — "please update our banking details" and "kindly amend the remittance account" are far apart lexically and close in embedding space.

3. Bodies and headers together. Headers are structured fields (display name, from, reply-to, subject, authentication results), not natural language. Two options: serialize them as text and concatenate with the body using separator tokens, or encode separately and fuse downstream. The word "unify" suggests the former. The value is cross-field inconsistency — a display name claiming one identity while the address and reply-to say another is only visible when both are in the same representation.

What does the term “Feature systems” mean? A set of features or something more complex?

Not "we have good features" but "we can define, compute, backfill, and serve features consistently at scale.

“Every email processed refines Abnormal's models, improving detection of novel attacks that legacy systems miss.”

Does “processed” here mean processed by Abnormal or by security teams? If it’s the former, how would that help future detection? It's not generating any labels just because it's processed

Unlabeled email adds nothing to a classifier. But "distribution of email data" is precisely the unlabeled half — and that does improve on volume alone:

  • Baseline sharpening. Every benign email tightens the estimate of normal for that sender, pair, and organization. A relationship with 200 observed messages supports a much sharper deviation estimate than one with 3.
  • Cold-start decay. New customers, new vendors, and new employees start with no history. Time and volume are the only cures.
  • Corpus for self-supervised pretraining. Masked-language-model training on business email needs no labels at all.

The novel-attack claim is the weaker link. Better baselines do help catch never-seen tactics, since deviation detection doesn't require a prior example. But converting deviation into a verdict is supervised, and that still needs labels — which arrive from customer reports and analyst adjudication, not from processing volume.

Figure from the post

  1. Is this screenshot for illustration or does it have any connection to how the alerts in front of security analysts actually appear?
  2. For the third highlighted indicator about the abnormal content and context, how is that being concretely defined as a feature?
  1. Illustration, not a product screenshot. That said, it's not pure fiction. The three categories almost certainly correspond to real feature groups, and Abnormal's actual case view does present an attack summary with contributing signals. So: accurate about what the system reasons over, invented in how it's rendered.
  2. Decomposing it into plausible actual features:
    1. Entity extraction from body text. The colored spans imply a tagger emitting typed entities — money amount, financial account reference, payment verb, date/urgency marker. These become categorical or count features: contains monetary amount, contains banking-detail change request, contains same-day deadline.
    2. Intent classification. A supervised head over BERT scoring the message against a taxonomy — payment redirection here. This is the "bypass normal process" part, and it's learned rather than rule-defined.
    3. Entity-versus-history comparison. "A bank never seen before" is not a text feature at all. It requires extracting the account or bank reference, then checking it against this vendor's previously observed payment details. That's a join between a text-extracted entity and the behavioral store.

What are the major ways Abnormal enables “Per-Customer Understanding“ in terms of features and modeling choices?

History systems capture each customer's behavioral communication patterns; from these they build a representation of what's normal for each user in that environment; detection then keys on fit to that environment rather than on attack indicators alone. No features or model structure are named.

Feature side (inferred) — three levels, non-overlapping:

  1. Relationship — per sender-recipient pair: prior message count, first-contact date, direction ratio, typical topics. Strongest single signal; first contact from a party claiming an established relationship is the classic anomaly.
  2. Identity — per sender: sending infrastructure (IPs, domains, auth results), display-name-to-address consistency, typical send hours, typical recipient set. Includes the vendor graph — which suppliers this customer actually works with, and what their normal payment details look like.
  3. Organization — who handles invoices, normal internal communication topology, what request types are routine here.

Modeling side (inferred) — the key choice is one global model, not per-customer models:

  • Features are relative, not absolute. Rather than raw "5 emails from this sender," the feature is deviation-encoded: z-score against this pair's history, days since first contact, distance from this sender's typical content. The customer-specific context lives in the feature value; the model weights are shared.
  • Why this and not per-customer models. Thousands of customers means thousands of training sets, most too small to fit. Deviation-encoded features let a global model learn what unusual-for-this-relationship means from pooled data across all customers, then apply it to any one. Small and new customers get the benefit of the whole population.
  • Consequence for cold start. A new customer has no history, so deviation features are undefined or noisy. The post-delivery API architecture allows backfilling from existing mailbox data — likely how the baseline is bootstrapped before live traffic accumulates.

The tension worth naming: the post claims per-customer understanding, but the durable version is per-customer features in a global model. That's the stronger design, and it's a different claim than what the marketing implies.

Figure from the post

How do tone and sentiment translate to actual features for prediction? On their own or compare average tone to current email’s tone?

Probably both.

Absolute features. Score the message alone: sentiment polarity, urgency, authority/pressure markers, formality. Useful because attack text has genuine population-level regularities — urgency and authority pressure appear disproportionately in payment fraud. Cheap, defined for first-contact mail where no history exists, and learnable from pooled labels across all customers.

Weakness: high false-positive rate. Urgent, terse, pressuring email is normal from plenty of legitimate executives and during quarter-end.

Comparative features. Score the message against this sender's or this pair's historical distribution: deviation in formality, in typical sentiment range, in usual topic set. This is where the signal actually lives — a CFO who has written 400 casual messages to this recipient suddenly sending a formal, urgent payment request is anomalous regardless of whether the text is alarming in isolation.

Comparative also directly serves the compromise case. After account takeover the address and authentication are legitimate, so the only remaining signal is that the writing doesn't match the person.

Concrete feature encodings (inferred — the post names none of this):

  • Distance between this message's embedding and the sender's historical centroid
  • Z-score of formality or urgency against this pair's history
  • Topic novelty: is this topic in the sender's usual set for this recipient
  • Cadence deviation, which the diagram lists separately: send time, frequency, and thread-timing versus the pair's norm

Why both must coexist: comparative features are undefined at first contact — no history, no baseline. Vendor fraud from a brand-new lookalike domain has zero prior messages, so absolute text features carry that case while comparative features carry the compromise case. Different attack types, different signal sources.

“Unlike most machine learning problems, this problem is adversarial.”

How specifically does Abnormal address the adversarial nature of the problem?

Model the defender's side, not the attacker's. The per-customer baseline claim is explicitly that they can flag deviation without anticipating what the attacker will conceal, because they understand the customer's environment better than the attacker does. That's the substantive adversarial argument — the attacker can freely mutate their own artifacts (domains, wording, templates) but cannot observe or mutate the victim's communication history. Attacking a baseline requires reconnaissance the attacker mostly doesn't have.

Inferred practices, none stated:

  • Prefer features the attacker can't manipulate. Sender infrastructure and message text are attacker-controlled; relationship history is not. Weighting toward the latter is the concrete version of the above.
  • Retraining cadence. Adversarial drift means distributions shift on the attacker's schedule, not seasonally. This raises the value of the feature-pipeline automation Shiebler described — time-to-ship-a-new-signal is the real defensive metric.
  • Ensemble diversity. Multiple models over different signal types raise the cost of finding an input that evades all of them at once.
2024 Paper original

PODTILE: Facilitating Podcast Episode Browsing with Auto-generated Chapters

Chapterization seems like it should enhance user experience. However, implementing it is tricky because podcast episodes are long and somewhat incohesive.

My notes

https://research.atspotify.com/2024/10/podtile-facilitating-podcast-episode-browsing-with-auto-generated-chapters

Summary

Context

Spotify has podcasts. A small % of podcasts have creator-provided chapters but the vast majority don’t.

Problem

Chapterization seems like it should enhance user experience. However, implementing it is tricky because podcast episodes are long and somewhat incohesive.

Solution

PODTILE, a fine-tuned encoder-decoder transformer to segment conversational data, simultaneously generating chapter transitions and titles for the input transcript.

Design Details

  • Architecture
    • Model:
      • Fine-tuned LongT5 pre-trained LLM
      • {Add more details about how it works}
    • Input:
      • Static context, dynamic context and transcript of episode
      • “Title: {title}, Description:{description}, Previous chapter titles: {previous chapter titles}, S1: {content of sentence 1} S2: {content of sentence 2} …”
    • Output:
      • “S1: {chapter title for chapter sentence 1}
  • Training
    • Data
      • Podcast dataset from a proprietary internal catalog: English episodes chapterized by their creators, contains 10.8k episodes, uniformly sampled with several filters. Split the resulting dataset into train/validation/test partitions of 8k/1k/1k episodes. Title and description used as the static context.
      • WikiSection - a Wikipedia-based dataset limited to two categories, en_disease and en_city, with normalized section titles for discriminative title prediction. They use only the English documents and use the title and abstract of each document as the static context.
      • QMSum - a collection of meeting transcripts annotated with topic segments and labels with 232 data points. They use the user-generated meeting summaries as the static context.
      • Format: Input chunks of up to 8000 words, with 7000 words dedicated to the document text and up to 1000 words to the metadata
    • Process
      • Batch size of 1, a learning rate of 5.0𝑒-5 (except learning rate 1.0𝑒-4 for Wikisection) with scheduler type of linear, and a maximum of 4 epochs.
      • Training on the podcast dataset took approximately 3 days. Inference of 1.1k episodes lasts an average of 1 hour.
  • Evaluation
    • Offline
      • Metric(s): WindowDiff, ROUGE (F1), SBERT (F1)
      • Data: training data explained in the previous section.
      • Baselines: CATS, Gen (seg + label) and GPT-4
      • Results: shown in Table 2 of Paper - PODTILE (either with only static context or with both static and dynamic context) was the best-performing approach
    • Online
      • Metric(s):
        • [For chapterization] Engagement ratios between episodes with auto-generated chapters and those with creator-provided chapters, # of chapter-initiated plays,
        • [As search component] nDCG, Recall (@30, 50 and 100) and RR
      • Results:
        • [For chapterization] Engagement ratio between 0.53 and 0.75 and 88.12% increase in chapter-initiated plays after the roll-out
        • [As search component] Results in Table 5
  • Deployment
    • N/A
  • Post-deployment
    • N/A

What I didn’t understand or am unsure about

Question

Answer

There are models that can take in audio input directly, which could potentially add additional signals like intonation, volume and speed compared to a purely text-based approach. Are these models likely to be useful for this sort of use-case or would they be largely redundant at best compared to the text transcription + language modeling approach in the paper?

Theoretically, it should have some value. Whether this value is sufficient to justify the investment is unclear. It might be worth trying with a small dataset and simple model to explore how well it works.

They do say at the end that “we aim to leverage other modalities, such as audio and video, to further improve chapterization”.

I have a hunch that the kinds of creators who provide chapter labels for their podcast episodes are not a representative sample of creators generally. Perhaps they’re more well-resourced or more diligent. This could lead to a model optimized for an unrepresentative type of content with misleading evaluation results. Is this a fair concern?

This is mitigated somewhat by the use of the other two datasets.

Perhaps some kind of analysis/breakdowns of the chaptered vs non-chaptered episodes would help identify any data distribution discrepancies.

“We augment the raw input text by adding index numbers before each sentence. This allows the decoder to predict the start of a chapter by referencing one of these indices.”

Does this mean that chapter breaks can’t be designated in the middle of sentences? I could imagine a situation where a speaker says something like “[content about topic 1] but let’s move on to [topic 2] and [content about topic 2]” where different parts of the sentence would belong in different chapters.

Yes, but I think it makes sense to do it this way; I think token/word-level indexing would make training slower/more expensive while probably not increasing performance significantly. It might be worth checking the transcript to see if there are any examples of a single sentence split into different chapters. If there aren’t, there’s even less of a reason to do token/word-level indexing.

Why use an encoder-decoder model instead of a decoder-only for this use-case?

The high-level reason is probably because the paper they based their work on used an encoder-decoder model.

An encoder-decoder model works by building bidirectional, context-aware hidden states from the input (encoder), and a decoder that generates the target sequence by querying those states via cross-attention.

As for why that paper used an encoder-decoder instead of decoder-only, let’s think about how a decoder-only approach would actually work. It could be one of two ways:

  1. Feed the entire transcript as a prefix into the model and then ask it to generate chapter boundaries. I think this is what they try with GPT-4.
  2. Feed sentence by sentence and ask it to determine after each sentence whether that’s the end of a chapter.

For approach 1:

  • The representation of each sentence is calculated by attending to all sentences before it. Sentence 100’s representation will not be influenced by Sentence 101’s representation.
  • Without this bidirectional conditioning at the representation level, the model's "understanding" of Sentence 100 is incomplete. It relies entirely on the decoder's ability to "look back" and compare disparate causal states, which is a more complex reasoning task than querying a bidirectional map.

For approach 2:

  • If you’re trying to evaluate if a sentence is a boundary based only on what came before it, you might trigger false positive boundaries at tangents or filler.

My guess is that contemporary state of the art decoder-only models might match encoder-decoder performance because there’s been much more investment in improving them.

Why did they use a batch size of 1 and 4 max epochs?

Maintaining the activations for a ~220M parameter model across such long sequences consumes massive amounts of Video RAM (VRAM). A batch size of 1 is typically the only way to fit such long-context windows into a single GPU's memory. Increasing the max epochs would slow down training.

In transfer learning for sequence-to-sequence tasks, models often reach peak performance within a few epochs (typically 3–5).

They don’t have a mechanism for ensuring that chapter titles are assigned sequentially, right? For example, it seems like they could have sentence 1 and sentence 3 be assigned to chapter A but sentence 2 be assigned to chapter B.

They provide chapter titles as dynamic context which the model could theoretically use to prevent this kind of ordering (if the training data also doesn’t have this property). Also, through supervised learning, the model can learn the statistical pattern that the next index generated (the number following the | separator) must be greater than the index generated before it.

Do they not use speaker identities in the transcripts? Figure 2 seems to suggest they don’t but why not?

Including speaker identities would create additional work during transcription and add additional tokens to the context window. I feel like it’s still worth at least exploring at some point though, maybe on a limited dataset to assess improvement potential.

Why 80/10/10 split?

They don’t provide an explanation but their dataset is large enough where 1000 test samples is probably enough to give you statistically significant results so you should leave the remainder for learning.

I doubt they put a ton of thought into this though since I don’t think it matters too much. 80/10/10 is pretty standard.

What’s the point of using the WikiSection dataset given that it’s a very different kind of format?

The use of WikiSection was specifically designed as a negative control. By showing that PODTILE didn't provide a massive advantage on short, structured text, the authors were able to prove that their technical contributions (Static and Dynamic context) were specifically solving the unique challenges of long-form, unstructured conversation.

Since the improvement on short, structured documents is smaller than longer ones, Spotify could disable the metadata-augmentation logic for short episodes if it seems costly or otherwise difficult to manage.

For future projects, it might also be helpful to know where PODTILE was helpful and where it wasn’t.

Did they fine-tune 3 separate versions of the base model for the 3 different datasets or do some kind of joint training?

I think they fine-tuned 3 separate versions, possibly because a single joint model would struggle with the inconsistent output vocabulary and stylistic requirements for the different datasets.

Why not use synthetic data?

This could look something like:

  1. Generate a general description of the podcast episode and its length.
  2. Generate plausible chapter titles.
  3. Fill in the chapters with text (potentially using a generative model fine-tuned on podcast transcripts)
  4. Evaluate your chapterization methodology on this dataset

There’s a possibility this could be helpful but we already have decent labeled real-world data so unclear if it’d be worth the cost. Might still be worth exploring at a small-scale or on heavily under-represented genres.

“CATS [52]: A multi-task learning model that combines boundary classification with coherent sequence detection”

What is the difference between boundary classification and coherent sequence detection?

Not super sure but I think coherent sequence detection tries to determine if the sentences in a sequence are logically ordered and semantically consistent. This training objective helps it develop a better sense of where topics naturally begin and end which can be useful for the boundary classification head.

“We use input chunks of up to 8000 words”

Why 8000?

PODTILE is built on LongT5-Base, which has a maximum sequence length of 16,384 tokens. In English, 1,000 words typically translate to roughly 1,300–1,400 tokens. Therefore, 8,000 words equal approximately 10,500 to 11,200 tokens. By capping the transcript chunk at 8,000 words, the researchers leave about 5,000 tokens of head-room. This space is strictly reserved for the Global Context (episode title and description) and the Dynamic Context (the list of previously generated chapter titles). Without this buffer, the model would be forced to truncate the very metadata that allows it to maintain a "global view" of the episode.

Also might be related to GPU memory constraints.

“Training on the podcast dataset took approximately 3 days”

Let’s say I want to make sure my training will actually do what I intend before I kick off this long and potentially expensive training process. What checks can I do beforehand to reassure me?

  1. Since you are using custom markers like S137, you must ensure the tokenizer treats them as atomic units. If the tokenizer splits S137 into S, 13, and 7, the model will struggle to learn the index logic. To check this, manually tokenize a few sentences and decode them. The output should show the marker as a single ID in the input_ids tensor. If it's split, you need to use tokenizer.add_tokens(["S137", ...]) and resize the model's embedding layer.
  2. Run a single training step with your maximum chunk size (8,000 words) and maximum output length. Use torch.cuda.memory_summary() or the PyTorch Profiler. Peak memory should stay under ~85-90% of your GPU capacity to leave room for the optimizer states and gradient accumulation.
  3. Overfit-to-One" Unit Test: Take a single episode and its ground-truth chapters. Train the model on only this sample for 50–100 iterations. The training loss should approach zero, and if you run inference on that same sample, the model should reproduce the ${index} := ${title} format perfectly. If it can't, your loss function or architectural wiring is broken.
  4. Run a "zero-shot" inference (greedy decoding) on the untrained model using your prompt: check that the model is generating tokens from the correct vocabulary (indices and words) rather than just repeating [PAD] tokens or producing infinite strings of punctuation.
  5. Gradient Health Monitoring: Log the Gradient Norm for the first 10–20 steps of training. The norm should be stable (e.g., between 0.1 and 5.0). If the norm is NaN or 0.0, your learning rate is too high/low, or your TGlobal attention is masking out too much information.

Explanation of WindowDiff

Imagine you have a long string of sentences. You take a fixed-size window (length $k$) and slide it across the text, one sentence at a time. At every stop, you count:

  1. How many boundaries are in the Ground Truth (Reference) window?
  2. How many boundaries are in the AI’s Prediction (Hypothesis) window?

Penalize any discrepancies.

What is BERTScore?

Instead of squashing a whole sentence into one vector, BERTScore breaks the process into four distinct steps:

  1. Contextual Embedding: Both the candidate and reference sentences are passed through a Transformer model (like BERT or RoBERTa). Every token gets its own unique vector that changes based on the words around it.
  2. Pairwise Cosine Similarity: It calculates the cosine similarity between every token in the candidate sentence and every token in the reference sentence. This creates a similarity matrix.
  3. Greedy Matching: For each word in the reference, it finds the "best match" (highest similarity) in the candidate sentence, and vice versa.
  4. Aggregation: These best-match scores are averaged to produce Precision, Recall, and F1 scores.

“We use SBERT title representations to apply soft-matching distance”

What do they mean by soft-matching distance?

Soft-matching distance is the technical solution to the "Assignment Problem": If your model predicts a different number of chapters than the ground truth, how do you know which predicted title should be compared to which reference title?

Use Sentence-BERT (SBERT) to perform a semantic "alignment" before calculating their final scores. Use Hungarian Algorithm to find the specific pairing of titles that minimizes the total soft-matching distance across the entire episode.

What is the difference between Matches_ref and Matches_pred?

These are trying to address “For every chapter the human wrote, how well did the AI cover it?” and “For every chapter the AI invented, how many of them actually correspond to a real topic?”

“We use SBERT instead of BERT for title representation because we measure distances

between entire chapter lists, where the atomic elements are titles. Unlike BERTScore,

which measures similarity between sentences at the word level for tasks like machine

translation and image captioning, SBERT is better suited for our purpose.”

I didn’t understand this.

The atomic unit for SBERT makes more sense for this task as we care about the title as a whole, not its constituent tokens.

“After examining a few examples, we speculate that lower performance in the dynamic context-only model may be due to a chapterization style different from the ground truth, hinting at the insufficiency of the state-of-the-art reference-based metrics and a single ground truth for chapterization.”

What would be a better way to do dynamic contextualization then? Could we add examples to the prompt or would chapterization heterogeneity mean that wouldn’t generalize?

Few-shot could help but that might really stretch our context window limits. Dynamic few-shot would be even better

Using reference-free metrics like human/LLM-as-a-judge evaluation might help.

“We compute the coefficient of variation … These results highlight the limitations of reference-based metrics used in Table 2 and show that dynamic context positively contributes to title quality, aligning with the original motivation for this feature”

I don’t get how the latter statement follows from the former.

I think it’s because dynamic context CoV was closer to the human PoV (inferred from Table 1) which implies that the dynamic context was able to mimic the rhythm of human chaptering better than the Static model.

“To test if longer static context enhances auto-generated chapter quality, we computed the Spearman rank correlation between static context length and the ΔROUGEL𝐹 1 of PODTILE with and without static context. We found a negligible negative correlation, suggesting that longer static context does not necessarily improve metric scores.”

What is negligible correlation? Does Spearman's rank correlation coefficient have any kind of statistical significance associated with it?

Yeah, so you can use p-values to determine if correlation is negligible.

“We compared engagement ratios

between episodes with auto-generated chapters and those with creator-provided chapters. A lower ratio would indicate that autogenerated chapters are less attractive or useful”

The engagement ratio metric is kind of confusing to me. How are we supposed to use this to determine if our model did well or not without any kind of baseline? Let’s say I get a value of 0.8 - is that good or bad, and how should that inform my decision making?

We do have the baseline of 1.0 but I still think it’s really subjective/arbitrary whether something like 0.8 is good and how we should make a decision based on that.

“Overall we saw an 88.12% increase in chapter-initiated plays after the roll-out.”

I’m assuming this is the total number of chapter-initiated plays, not the rate per podcast with chapters?

Yes, I think so, which makes me question its utility as a metric. If we have more episodes with chapters, chapter-initiated plays are bound to go up.

Without any better alternatives though, might as well put the number there (although contextualization would have been nice).

“In contrast, auto-generated chapters have a more balanced distribution, with 50.84% of play counts from super light to upper medium users. This shows auto-generated chapters help users with limited time navigate episodes efficiently”

The latter statement is only true if we’re speaking relatively, right? There is no objective quantification of its utility. All we’re saying is that it was better for these kinds of users than it was for other kinds of users.

Yeah, I think this phrasing is imprecise but also doesn’t really matter too much.

“However, in entertainment categories like ‘TV and Shows’, ‘Leisure’, and ‘Arts’, longer episodes receive more chapter plays, suggesting that both duration and content influence chapter usage.”

This low-key feels like p-hacking. How to avoid multiple testing problems when doing these sorts of breakdowns?

We can use various multiple testing corrections but this feels overkill - this statement is not really a core part of their paper so directional signal is probably fine.

“This approach could reduce costs by at least tenfold compared to indexing entire transcripts.”

How much does it actually cost to index entire transcripts? A tenfold cost reductions mean a lot when the baseline is $x million but not when it's $x thousand.

My guess is it’s closer to a million than a thousand, based on the size of the data (episodes x tokens) so I think this it’s fair to cite this as a major benefit.

Would the indexing cost still be really high if they used semantic search instead of BM25?

Probably even higher since dense vectors take more data.

Is RR in Table 5 MRR (Mean Reciprocal Rank) or something else?

Yes, I think that’s what it is.

I have a feeling that chapterization improves user experience slightly in ways that are difficult to capture with conventional metrics. What would be a good way to actually evaluate this?

I’m guessing A/B testing for retention, churn or engagement would require a prohibitively large sample size since I don’t think the effect would be that large. User surveys, perhaps?

Alternatives include

  1. Measuring "Friction"
    1. Compare a cohort with chapters vs. a cohort without.
    2. If users in the chapter group reach a "Target Segment" (determined by their search query) with fewer back-and-forth scrub actions, you have objective proof that you’ve reduced cognitive load. You’re looking for a reduction in "random seeking" behavior.
  2. Pairwise Preference
    1. Show users two versions of the same episode. Version A has PODTILE chapters; Version B has human chapters (or no chapters). Ask them "Which of these makes you feel more confident that you know what's in this episode?"
  3. Task Success Rate
    1. Give users a specific question (e.g., "What did the guest say about the 2008 financial crisis?") and a 60-minute audio file.
    2. Measure Time-to-Success and "Perceived Ease of Use" (using a NASA-TLX or SUS scale).
  4. LLM-as-a-Judge
    1. Prompt: “I am a busy professional. I only care about [Topic X]. Rate how helpful these chapter titles are for skipping the fluff."
    2. Use a 5-point Likert scale specifically for Informativeness and Boundary Logic.

"After deploying the model on our platform, we observed that users find auto-generated chapters helpful for browsing episode content.”

Engagement doesn’t imply helpfulness though?

Yeah, I think this phrasing is imprecise. Might be worth tracking engagement over time to account for novelty effects.

Other thing worth tracking

  • If users click and then immediately scrub manually anyway, the chapter wasn't helpful—it was just a failed starting point.
  • Bounce Rate per Chapter: How many users click a chapter and leave the app within 10 seconds?
  • Feature Stickiness: Does a user who uses a chapter on Monday come back and use a chapter on Friday?

“Therefore, we plan to extend our evaluation to include reference-free metrics.”

What would reference-free metrics look like here?

  1. SBERT cosine similarity between generated chapter title and actual transcript text of segment
  2. LLM-as-a-Judge:
    1. Informativeness: "On a scale of 1-5, how much does this title reduce the user's uncertainty about what happens in this segment?"
    2. Brevity/Punchiness: "Is the title concise enough for a mobile screen (under 40 characters) while remaining descriptive?"
    3. Boundary Logic: "Does the chapter start exactly when the topic changes, or is there 'leaked' content from the previous topic?"

How would I convince a business leader that this project was worth doing?

I don’t think any of the offline or online metrics for the chapterization use-case could help make a really compelling case.

The search improvements on the other hand seem useful; increased search relevance -> better user experience -> higher retention/less churn, positive word-of-mouth, increased willingness to upgrade from free plan to paid plan.

How would you improve this system if you were building it today?

  1. Multimodal
    1. Audio and speaker identities
  2. Models today have 1M+ context windows; we can use "Global-Local" Attention, feeding the entire transcript into a model at once
  3. Two-pass workflow
    1. Check for hallucinations and edit
    2. To save resources, maybe only do this for high-engagement podcast episodes?
  4. Direct Preference Optimization (DPO) or RLHF based on actual Spotify engagement data to align the model with User Utility rather than Human Mimicry.
  5. Classifier that detects the Podcast Genre; then modulate chapter density

Level of understanding: 🙂

2024 Blogpost original

How DoorDash leverages LLMs for better search retrieval

Traditional query segmentation methods often struggle to capture meaningful word segments for more complex queries. String-matching has limited recall due to its relatively narrow search capacity (no semantic understanding).

My notes

Summary

Context

Users can search for items or stores/restaurants on DoorDash. DoorDash has built knowledge graphs for both food items and retail product items. These graphs define relationships between different entities.

User queries can be mapped to product attributes present in the knowledge graph using a two-step process: query segmentation methods like n-gram analysis and mapping these generated segments to concepts available in our knowledge graph using string-matching.

Problem

Traditional query segmentation methods often struggle to capture meaningful word segments for more complex queries.

String-matching has limited recall due to its relatively narrow search capacity (no semantic understanding).

Solution

Use an LLM to perform query segmentation and entity linking; the extracted attributes can be used as filters or signals for search retrieval/ranking.

Design Details

  1. Data
    1. Knowledge Graph
    2. User Search Data
  2. Architecture
    1. Query Segmentation: Use an LLM to identify meaningful segments and categorize them under DoorDash’s taxonomies
    2. Entity Linking: Augment the LLM’s context using the closest 100 taxonomy concepts (as determined by query <> entity embedding similarity) for each query and prompt the LLM to link queries to corresponding entities from specific taxonomies such as dish types, dietary preferences, cuisines, etc
  3. Training
    1. N/A
  4. Evaluation
    1. Offline:
      1. N/A
    2. Online:
      1. Metrics: Whole Page Relevance, engagement and conversion.
      2. Results: {TO ADD}
  5. Deployment
    1. N/A
  6. Post-deployment
    1. {Retraining?}

What I didn’t understand or am unsure about

Question

Answer

“Even though the hallucination rate on the segmentation process is low, … we also benefit from the immediate classification of the output in a valuable category for our retrieval system”

What does this mean?

I think they’re saying that the combination of high accuracy + taxonomic output structure together make the LLM well-suited for initiating the retrieval process.

Do they require that every piece of the query be segmented into a taxonomy or only those that can be easily fit into those categories?

Probably not. If they required every word to be categorized, the system would become brittle and prone to hallucinating categories for words that don't belong in one.

Their system can’t handle stuff like “cheese pizza or pepperoni pizza” easily, right?

Doesn’t seem like it could but I guess you could modify the LLM’s prompt to handle this and pass down two processed queries to the next layer. Worth exploring the prevalence of search queries with this sort of request first though.

For the query segmentation part, is the structured output restriction part of the prompt or enforced in some other way during inference?

On a related note, is there an easy way to enforce the structured output restriction when we have many different possible taxonomies? What would that actually look like concretely?

They probably included all taxonomies in the prompt. For the Constrained Decoding part, you can define a JSON scheme that would look something like this (note that all fields are optional):

{

"type": "object",

"properties": {

"Dietary_Preference": { "type": "string" },

"Flavor": { "type": "string" },

"Product_Category": { "type": "string" },

"Quantity": { "type": "string" }

}

}

Later on, they also say “To ensure [high precision], we developed post-processing steps to prevent potential hallucinations in the final output and ensure the validity of both our segmented queries and their linked entities”

“Because the knowledge graph has been ingested into the search index as part of our document understanding work, we can make many rich attributes available for retrieval”

What “does ingested into the search index” mean exactly?

Inverted Index with list of attributes containing restaurant/item IDs.

Why 100 taxonomy concepts as opposed to any other number or a similarity amount threshold? How would you evaluate the effect of different options?

An amount threshold (e.g. 0.7) on its own could lead to a very wide range of results returned which in extreme cases could impair the pipeline’s performance. I think they could still use a hybrid approach though like picking max(Top 100, min(items with similarity > 0.7, 200)).

Evaluate using Recall@K (as labeled by humans or an LLM judge), Precision@K and maybe some latency/cost related metrics. Decide how you want to weigh each of these and pick accordingly.

Why do they filter for “Dish Type: Chicken Wings” first and then for “Taste:Spicy”? Is this just for illustration or is there any deeper reasoning behind that ordering?

Not fully clear from the blogpost but in theory, the team could determine which taxonomies are filtered for first based on what seems like it’d be most useful.

“This process ultimately generates a set of linked taxonomy concepts for each query that we can use directly to retrieve items from the search index.”

Is this just retrieval basically just filtering for items that match each attribute in the index?

For retrieval, yes but in the ranking stage, attributes might be used more holistically (also depends on the MUST vs SHOULD part).

“This makes it easier to control what to retrieve by implementing a specific retrieval logic, such as making all dietary restrictions a MUST condition and allowing flexibility of less strict attributes such as flavors as a SHOULD condition.”

Who decides what’s a MUST and what’s a SHOULD?

My guess is that this is manually determined by the DoorDash team based on their intuition for how users search. You could also A/B test different approaches and see what works best.

“Annotators review a statistically significant sample of the output to verify that query segments are correctly identified and accurately linked to the appropriate entities in the knowledge graph.”

What would be the right metric for such an evaluation? Also, how would you decide the sample size needed for statistical significance without knowing what the results will look like?

  • Precision (Attribute Accuracy): The percentage of identified segments and links that are correct. (e.g., If the model says "vegan," is it actually vegan?)
  • Recall (Segmentation Coverage): The percentage of meaningful segments in the raw query that the model successfully captured. (e.g., Did the model miss the word "spicy"?)
  • Exact Match (EM) Rate: The percentage of queries where the entire structured JSON output (all segments and all links) is 100% correct.
  • Link Accuracy: A specific sub-metric for the Entity Linking stage: given a correct segment (e.g., "no-milk"), how often did the model choose the correct official taxonomy concept (e.g., "Dairy-Free")?

You can use a conservative estimation approach, assuming maximum variability (p = 0.5); then decide some margin of error like 5% and a confidence level like 95%.

“Feature staleness: Some segmentations and links likely become stale over time.“

Why would they become stale?

  1. Knowledge graph updates
  2. User linguistic drift (existing users start using different kinds of terms to search or different kinds of users with different linguistic preferences start using the platform)
  3. Seasonal dynamics

Are they doing any kind of semantic caching or just exact query matches?

“Using LLMs for batch inference on a fixed set of queries can provide highly accurate results” implies just exact query matches. Semantic caching has some risk of worsening precision but I feel like it’s still worth testing at least.

They do have non-LLM approaches based on embedding retrieval and BM25 which means cache misses may not be a huge deal.

“A hybrid approach strikes the right balance between memorization and generalization.”

What would such a hybrid approach actually look like? Reciprocal Rank Fusion (RRF)?

Since they talk about “[retraining] our ranker with a more comprehensive dataset”, the ranker might actually be an ML model like a GBDT where the LLM-derived attributes act as features.

“As the rankers caught up…”

What do they mean by caught up here?

  1. Feature integration: enabling the ranker to accept LLM-derived attributes as features.
  2. Training the model on an updated dataset including those features.

“We saw a substantial increase in the trigger rate of popular dish carousels ...”

What does trigger rate mean here? % of searches where the dish carousel appeared?

Yeah. The higher trigger rate might be a result of improved retrieval which, when coupled with minimum item # requirements for triggering the carousels, would increase the trigger rate.

“we observed nearly a 30% increase over our baseline, which also means we are aligning search results more closely with consumer intent, making it easier for them to place orders”

Why does a higher trigger rate for dish carousels on its own indicate search results more closely aligned with consumer intent? Or is this statement based on the evidence presented further down?

They’re treating trigger rate as a proxy for recall; the validity of this depends on the criteria for a carousel being triggered.

How is whole page relevance defined and actually measured?

Maybe nDCG (normalized Discounted Cumulative Gain)?

To actually measure this, you can have human Search Evaluators rank each result as relevant or not relevant and then calculate nDCG.

“Helping users rewrite queries and recommending search paths they can explore”

From a UX POV, would this just show up as search recommendations?

Also, how exactly would you go from a user query to these recommendations?

These could appear as search recommendations (options below the query before the user presses enter) or as “Did you mean” after the user presses enter.

I was thinking something along the lines of

User query -> query optimized for retrieval by LLM by segmentation and entity linking -> present optimized query to user to check if that's what they mean

The tricky thing here would be that the retrieval-optimized query is no longer a human-understandable phrase but a mapping of attributes. I don’t know if there is a good + fast way to reverse engineer a query out of that.

“Showing new users which queries they may want to search”

How is this different from the previous point?

I think this is for the empty search bar state.

“Deeper query and catalog understanding let us better understand the overlap of attributes between entities and create personalization signals”

What does this mean concretely?

Moving away from personalizing based on what you bought (e.g., "You bought pizza, here is more pizza") and toward personalizing based on the DNA of what you like (e.g., "You like spicy, high-protein, late-night meals").

Because of catalog understanding, that pizza is now a bundle of attributes: (Topping: Pepperoni, Flavor: Salty, Texture: Cheesy, Prep_Time: <20min).

These attributes can then be coupled with a user’s historical search/engagement data to personalize results.

This is pretty helpful for tackling data sparsity too. Instead of just knowing you like pizza, the system knows what specific attributes you like so each individual search contributes much more to improving future personalization.

How would I convince a business leader that this project was worth doing?

“Online testing also showed that increased relevance aligns with an increase in engagement and conversion.”

Higher engagement and conversion means more $ value directly and indirectly through improved retention, positive word-of-mouth, etc.

How would you improve this system if you were building it today?

  1. Multimodal catalog ingestion
    1. For example, even if a dish isn’t tagged as spicy, you might be able to infer that it’s spicy from a picture
  2. Agent chooses MUST vs SHOULD instead of using pre-determined strategies
  3. More dynamic personalization
    1. The results you see should include additional real-time personalization updates so that your current search session behavior informs what results you see in subsequent searches within the same session.

Level of understanding: 🤓

2025 Paper original

Contextualizing Spotify’s Audiobook List Recommendations with Descriptive Shelves

Spotify wants to develop a system that would recommend a shelf of contents a user will likely engage with and package the shelf with an appropriate title that describes the contents and provides context. However, content-based explanations…

My notes

Summary

Context

Spotify has audiobooks. The standard audiobook recommendations consist of a section called “Audiobooks for you” generated through a graph neural network algorithm.

Problem

Spotify wants to develop a system that would recommend a shelf of contents a user will likely engage with and package the shelf with an appropriate title that describes the contents and provides context.

However, content-based explanations for recommendations often rely on extracting information from user reviews, user-generated tags, or using items’ rich metadata which aren’t always available.

Solution

To overcome this cold-start problem, they use LLMs to enrich the content of the items in the catalog with descriptors that are grounded on item metadata. Based on the enriched item metadata, the pipeline generates descriptive shelves, which group recommendations into thematic and personalized lists.

Some key insights/new things I learned

  1. This seems like a solid way to provide explainable recommendations without having to generate explanations for every user (which would be very costly)

Design Details

  • Data
    • Taxonomy for LLM developed using internal audiobook search query data and requests issued at Reddit on the /r/booksuggestions/ forum
    • Interaction data to generate recommendations from this paper
    • Internal catalog data for audiobooks with title, author(s), description, and BISAC genres.
  • Architecture
    • For each audiobook in their catalog, use an LLM to extract the 10 types of descriptors in the taxonomy
    • For each user, generate a candidate set of recommendations using model from this paper
    • Create a shelf for each unique descriptor associated with the candidate set
    • Diversify the list by removing descriptors that are too similar to each other
    • Populate each shelf with items tagged with the corresponding descriptor, and then rank each item according to the recommender system scores for the user
  • Evaluation
    • Online Test 1: Audiobooks for you vs descriptive shelf (A/B test)
      • Metrics: “discovery metrics” and “engagement metrics”
      • Data: real-world user interactions in A/B test within Spotify homefeed
      • Results: increased discovery metrics but decreased engagement metrics.
    • Online Test 2: Automatically generated descriptive shelf vs editor-curated descriptive shelf (A/B test)
      • Metrics: impression to click rate, impression to stream rate, # impressed, # interacted
      • Data: real-world user interactions in A/B test from Audiobooks subtab within Spotify homefeed
      • Results: increased discovery metrics and engagement metrics.
  • Deployment
    • N/A
  • Post-deployment
    • N/A

What I didn’t understand or am unsure about

Question

Answer

“To define such taxonomy, we looked into … requests issued at Reddit on the /r/booksuggestions/ forum”

Is this data obtained through an API or web-scraping? I heard about policy changes at Reddit that would restrict use of the platform’s content so curious whether it applies here.

Depends on the volume. Seems like the Reddit API remains free for low-volume academic/research use under specific tiers.

The different elements of the taxonomy don’t seem mutually exclusive - for example, Juvenile Fiction and Children’s Literature. Is taxonomy overlap something to be concerned about?

The authors address overlap to some extent during the Descriptor Ranking and Diversification phase (Section 3.2.2).

What LLM would make sense to use for the audiobook taxonomy mapping?

At Spotify’s scale, my guess is that hosting an open-source model via vLLM (or similar alternatives) is probably cheaper than proprietary APIs. With that in mind, DeepSeek-V3 seems to be the top open-source model these days although at the time of writing, I think Llama 3 was best.

We could also use teacher-student distillation in which case our final model would be a smaller model than the ones mentioned. If we have labeled data, we can test out these different approaches and see how it goes.

Note that we might also benefit from using constrained outputs to ensure the taxonomy is correctly generated.

“The metadata used as input to the LLM is the title, author(s), description, and BISAC genres”

Why not use the plot of the book as input for the LLM? Something like the Enemy to Lovers story trope seems like it would likely only be inferable from the plot (like that on a book’s Wikipedia page) as opposed to a short description or the other metadata.

The more data you put into the LLM prompt, the slower it is.

They probably also just didn’t have access to that data in an easy-to-use form.

“For each audiobook in the catalog, we use the LLM to extract the 10 types of descriptors.”

This process seems like it would take a while, especially with a long prompt and large audiobook database like it seems they had. What are good ways to make this faster?

  1. Automatic Prefix Caching (APC): Since the instructions and the 10-dimension taxonomy remain identical for every book, the model can cache the "keys" and "values" (KV-cache) for that specific part of the prompt. Instead of re-reading the long prompt millions of times, the LLM only "reads" the new book metadata, reducing the Time to First Token (TTFT)
  2. Constrained Decoding (Structured Output). This prevents the model from generating "chatty" filler (e.g., "Here are the tags for this book...") which wastes time and compute power.
  3. Quantization
  4. Speculative Decoding
  5. Model distillation: use the teacher to label a golden set and train a student to mimic the teacher’s labels. I’m thinking something along the lines of this blogpost.

“Manual and automatic evaluations showed high accuracy in the task”

What would an automatic evaluation of this look like?

I’m assuming they have labels from human review.

Classification metrics against a gold standard

  1. Compare generated taxonomy mapping to labels
  2. If each taxonomy component can only have 1 label, we can compute aggregated accuracy across all the taxonomy components
    1. For example, the LLM generated (1) Genre: Juvenile Fiction, (2) Theme: global politics, (3) Characters Descriptions: Female Protagonist, and the human generated (1) Genre: Juvenile Fiction, (2) Theme: American politics, (3) Characters Descriptions: Female Protagonist
  3. If we can have more than one label (for example, Genre: Juvenile Fiction, Romance), we can use Macro-Precision/Recall instead.
  4. Note that semantic similarity would not be considered here, which makes this approach suboptimal

LLM as a judge

  1. Use an LLM to compare the taxonomies and generate the same metrics as above but ask it to score matches on a scale from 0 to 1, instead of binary, to account for partial matches like “American Politics” vs “North American Politics” and synonyms.
    1. We might also use a smaller language model like BERT to generate embeddings for generated taxonomy components which can then be used to compare similarities using cosine similarity.

"from a set of top-K recommendations and their respective descriptors, we … diversify the list by removing descriptors that are too similar to each other"

What exactly would the process of deduplication look like here? It seems like they’re using content embeddings but what model makes the most sense here and how do you set a threshold for filtering?

Sentence-BERT probably makes the most sens; you can experiment with different threshold values to see what seems better although I don’t know the exact quantitative approach you would take for this.

“we rank and filter the set of candidate items for each of the distinct descriptors”

Does filtering just mean that the candidate item didn’t have that descriptor? Could this lead to shelves of very different sizes (which might be a suboptimal user experience)?

You could set an allowable range for these shelves; any shelf with more than the max allowed items will have its lowest ranked items pruned. Any shelf with fewer than the least allowed items will either not be shown or (less likely) have its filters relaxed to allow it to include other related content.

“we first use a ranking function to sort the set of distinct descriptors that characterize the items in l according to their relevance to the user”

What does this mean?

Maybe they have some way to see how relevant a descriptor is to a user? This could be by looking at what % of their recommendations or historical engaged items were associated with that descriptor.

{To investigate further}

“If we use more fine-grained and specific

descriptors for shelves, we end up with narrow descriptors that might be overly specific for a user, whereas coarse-grained descriptors might not reveal enough about the audiobooks in a shelf and encompass a large part of the catalog.”

I understand why this is an issue but do they explain how they address it?

They briefly mention combination descriptors that could theoretically help narrow down overly broad descriptors although they don’t cite it as a way to specifically target this issue.

On the flip side, shelves generated from overly narrow descriptors can just be discarded or have their filters relaxed to allow them to include other related content.

"The approach we use is to remove items that do not have descriptors matching the chosen descriptors and rank them according to the recommender system scores for the user."

Would it be helpful to also include some additional relevance scoring between the descriptor and the item as part of the ranking?

I think it could be helpful but it’d probably increase latency by some amount. Worth experimenting to see what happens.

What is the difference between # impressed and # interacted?

# impressed means how many people saw it; # interacted is how many people engaged with it.

I’m not entirely sure why # impressed is a useful metric but {will think more about it}.

In the second A/B test, they compare their approach with the manually curated shelves.

I’m assuming these manually curated shelves are not personalized the same way which feels like an unfair disadvantage.

Also, why do they no longer compare to the existing “Audiobooks for you” shelf? Does this shelf not appear when you open the audiobook subtab from the homepage?

It is a disadvantage but that’s sort of the point; an automated system is capable of personalization at a much larger scale so it’s worth knowing if the improvements from personalization can compensate for any deterioration from lack of human involvement in curation.

I’m not super sure about the second part but I think it might be because they’re considering two possibilities:

  1. The regular homepage contains “Audiobooks for you” which they’re considering replacing with the descriptive shelves.
  2. The audiobook subtab accessible from the home page historically contained static, manually-created shelves that they’re proposing replacing with this automated process.

How would I convince a business leader that this project was worth doing?

Context: Spotify main revenue sources are ads and user subscriptions. Costs are much more varied.

Assuming that the audiobook subtab actually used to contain manually-created shelves, the metrics in Table 1 should be fairly compelling on their own.

To further translate this to business lingo, we can highlight cost-savings from reducing/eliminating the editorial curation, and connect increased user-engagement and discovery metrics to improved retention and positive word-of-mouth (which should increase revenue from both major streams).

How would you improve this system if you were building it today?

  1. Consider the audiobook format more comprehensively
    1. Multimodal to ingest audio and generate taxonomy mapping
    2. Acoustic embeddings
  2. Adjust Top-K then Filter
    1. Instead of filtering a pre-ranked list, you search the entire catalog for the best matches for a theme (e.g., "Small Town Mystery") that also align with the user's vector
  3. Dynamic Taxonomy Scaling
    1. For New Users, the system should default to high-level, high-confidence "Coarse" descriptors. For Power Users, the system should trigger "Fine-grained" descriptors (e.g., "Hard Magic Systems" or "Enemies-to-Lovers Tropes").

Level of understanding: 🤓

2024 Paper original

OmniSearchSage: Multi-Task Multi-Entity Embeddings for Pinterest Search

Search relevance and engagement could be better.

My notes

Summary

Context

Users can search for stuff on Pinterest: these searches can surface Pins, products and suggested search queries.

Problem

Search relevance and engagement could be better.

Solution

Unified query, pin and product embeddings using enriched entity representations.

Design Details

  • Data
    • {TO ADD}
  • Architecture
    • {TO ADD}
  • Evaluation
    • Metric: {TO ADD}
    • Data: {TO ADD}
    • Results: {TO ADD}
  • Deployment
    • {TO ADD}
  • Post-deployment
    • {TO ADD}

What I didn’t understand or am unsure about

Question

Answer

“Pinterest’s search feed encompasses a diverse range of content”

Search feed here just means the search results, right? It doesn’t have any connection to the main feed

Yes.

“Embeddings can power retrieval use cases via approximate nearest neighbor (ANN) search [and] enable detailed content and query understanding in ranking models without the

overhead of processing raw data”

Retrieval is first-round filtering and ranking is done to reorganize those first-round retrieved results according to a more sophisticated scoring mechanism. When they say embeddings are used for retrieval and ranking, is it generally true that bi-encoder embeddings are the main kinds of embeddings used for retrieval and cross-encoder embeddings are the main ones used for ranking? Or are there other kinds of embeddings used in retrieval and ranking?

Bi-encoder embeddings are indeed the main kinds of embeddings used for retrieval.

For ranking, cross-encoders are common but we can also use pre-computed embeddings as high-level input features for the ranking model, rather than the ranking model necessarily being a raw cross-encoder. This final model could be a GBDT or a deep neural network.

Graph-based embeddings (which in this case could be derived from Pinterest’s user-pin engagement graph) can also be used for retrieval or ranking.

“Embeddings can … serve as a strong base to learn in low-data use-cases”

What does this mean?

It means embeddings can be used for transfer learning. Embeddings trained for one situation can meaningfully capture some underlying structure of the content or entity that is useful for other use-cases.

When they say product, do they just mean an ad listing? Later on, they say “[OmniSearchSage] powers embedding-based retrieval for standard and product pins, queries and ads” so I’m confused what a product is supposed to be.

Product Pins are Pins that specifically represent a buyable item, containing metadata such as price, availability, and a direct link to a checkout page.

Standard Pins are regular user-curated Pins (images or videos) intended for general discovery rather than immediate e-commerce.

Ads can be standard Pins or product Pins.

“A single query embedding can be used to retrieve queries, products, and Pins with nearly the same effectiveness as task-specific embeddings.”

When they say task-specific embeddings, do they mean embeddings that were trained separately on <query, query>, <query, product> and <query, Pin> pairs as opposed to the training data being combined?

Yes.

“A single query embedding can learn compatibility with multiple pre-existing embeddings and learned entity embeddings”

What does it mean for a query embedding to “learn compatibility”?

In the context of the paper, "learning compatibility" means training the query encoder so that the vectors it produces are mathematically aligned with vectors from other models in a shared embedding space.

The model is trained using a multi-task objective (contrastive loss) that forces query embeddings to be "close" (high cosine similarity) to relevant items (Pins, Products) and "far" from irrelevant ones.

OmniSearchSage doesn't just learn new representations; it learns to map queries into the same numerical coordinate system used by pre-existing models like PinSage or ItemSage.

Do two tower models for search retrieval generally work the same way as they do for RecSys?

My understanding is that the two towers in RecSys are Users and Items. In search, it seems we replace the user tower with a query tower - does that mean RecSys-like personalization isn’t possible?

RecSys prioritizes long-term user interest whereas search systems focus on immediate user intent.

Personalization can still be done by adding features like location in the query.

Personalization is also usually done in the ranking stage so we don’t need to do it during retrieval.

“Early-stage fusion can pose computational hurdles”

Why?

In the context of this paper, early-stage fusion refers to the way multimodal features (like text and images) are combined within a single entity (e.g., a Pin) before the final embedding is generated.

Early fusion typically requires a complex model (like a Transformer or Cross-Attention layer) to process the raw interactions between different modalities (pixels vs. text tokens) simultaneously. This is computationally expensive to run for billions of items in real-time.

The paper is referring to the Entity Encoder itself. If you use a heavy "early fusion" model to generate the Pin's embedding from its image and text, you create a massive offline indexing bottleneck and high online latency for any entity that needs to be encoded on the fly.

The authors opted for a late-fusion approach for the entity encoder to maintain efficiency. They process different modalities (text, visual, graph) through separate specialized encoders (like DistilBERT and ItemSage) and then combine their high-level embeddings at the very end. This modularity allows them to "swap" or update individual modality encoders without redesigning the entire interaction layer.

Brief summary of PinSage, assuming no familiarity with GNNs

On Pinterest, users save Pins to "Boards". If a Pin of a mid-century chair and a Pin of a vintage rug are often saved to the same "Living Room Ideas" boards, the system learns they are contextually related, even if they look completely different.

A Pin "chats" with the other Pins it shares boards with. It gathers their visual and text information to update its own "identity". It repeats this process multiple times. This means it doesn't just learn from its direct board-mates, but also from its board-mates' board-mates, capturing very subtle relationships across the entire platform.

Why are PinSage and ItemSage different?

From a brief glimpse at the ItemSage paper, it seems they tried a graph-based approach as well for products but it didn’t work great.

{Will add more detail once I’ve fully read through the PinSage and ItemSage papers}

“OmniSearchSage … offers a unified query embedding model that jointly trains query embeddings for query-query, query-pin, and query-product retrieval and ranking”

They mention retrieval and ranking together here. If they’re using OmniSearchSage for ranking, does that mean they’re not using a cross-encoder? Or would that be an additional stage (re-ranking?)

Probably not using a cross-encoder, just using bi-encoder embeddings produced by OmniSearchSage as input features for a downstream ranking model. They’re serving 300k requests per second so it makes sense to try to reduce latency.

“As a side effect, we also get compatibility with some embeddings due to the triangle inequality property inherent to cosine similarity”

What does this mean?

If the model learns that Query (Q) is very similar to Pin (A), and Pin (A) was already designed (by an older model like PinSage) to be very similar to Pin (B), then mathematically, Query (Q) must also be somewhat similar to Pin (B).

Why did they use BLIP instead of CLIP?

CLIP is a purely contrastive model and isn’t directly capable of generating captions; it needs to be combined with a decoder.

“For a robust assessment, three distinct ratings were collected for each image within a sample of 10𝑘 images, curated uniformly across various broad pin categories.”

Let’s say I was working on this project and had to determine the number of images to get labeled for this assessment. How would I do that?

{I’m guessing this number was chosen at least somewhat arbitrarily but if not, the statistics are a bit challenging to think about so I’ll look into this later - Cohen’s kappa}

“We exploit this user-curated information by accumulating the titles of all boards each pin has been saved to.”

Wouldn’t it make sense to do some sort of clustering on these titles to try to capture distinct boards?

I'm thinking more about an example they’re only getting the Top 10 board titles and the Top 50 looked something like this:

Position 1: Unique Title 1

Position 2: Unique Title 2

….

Position 10: Unique Title 10

Positions 11-50: Slight variations of Title 11.

Title 11 and its variations would be excluded entirely even though they are more relevant.

I think clustering would be an improvement but I don’t know if it’d be enough of an improvement to justify the effort since the example presented seems like a fairly unlikely situation.

“First, each title is assigned a score influenced

by two factors: its frequency of occurrence and the prevalence of its comprising words”

What does “the prevalence of its comprising words” mean?

I think "prevalence of its comprising words" refers to how common or "generic" the specific words in a board title are across the entire Pinterest platform.

By accounting for prevalence, the system can deprioritize common, low-information words and prioritize titles that are "specific and descriptive" for that particular Pin.

Explanation of Figure 1: Diagrammatic Representation of OmniSearchSage’s Multi-Entity, Multi-Task Architecture.

  1. What does L(...) represent?
  2. Why do we have 2 models each for the item and the Pin? What do the unified encoders represent?

L is loss. Each line ending in an L corresponds to a specific retrieval objective. L(query, pin} is the loss for matching a query to a Pin using the new unified encoder, while L(query, pin_c) is the loss for matching a query to a Pin using the pre-existing PinSage model (the "c" stands for compatibility). The model jointly optimizes the sum of all these losses.

There are two models for each because Pinterest is optimizing for two different goals at the same time: Future Performance and Backward Compatibility. The unified encoder is trained from scratch. Its purpose is to learn the best possible new representation for Pins and products by using updated features like LLM captions and board titles.

By including the legacy models in the training, the system forces the new Query Encoder to produce embeddings that work with both the new system and the old systems already running in production. This allows for a "seamless migration" without having to re-index the entire Pinterest catalog of billions of items overnight.

Brief summary of DistilBERT, assuming some familiarity with transformers and BERT.

DistilBERT used BERT as a teacher to train a smaller student model to mimic its output probability distributions. Roughly 97% of BERT's performance while being 40% smaller and 60% faster

“Overview of the query encoder architecture. The encoder takes the output from the last layer associated with the ‘CLS’ token”

Why do they take the output from the last layer associated with the CLS token?

In BERT-based architectures (like DistilBERT), the [CLS] token is a special "Classification" token prepended to the start of every input. Through the self-attention mechanism, the model is trained so that the final hidden state of this specific token captures a consolidated semantic representation of the entire sequence.

Using the [CLS] token is computationally more efficient than alternatives like "mean pooling" (averaging the vectors of all words in the query), as it requires extracting only a single vector from the last layer rather than performing an additional mathematical operation across all token outputs.

“In our model, we utilize a single unified encoder for both pins and products”

How does this relate to the ItemSage and PinSage they mentioned earlier? Why do we have this additional encoder?

The Unified Pin and Product Encoder does not replace ItemSage and PinSage; rather, it incorporates them as input features to create a more comprehensive representation.

In the OmniSearchSage architecture, PinSage and ItemSage are treated as "pretrained and frozen" continuous features. The Unified Encoder (shown in Figure 3) takes these pre-existing graph embeddings and combines them with other raw data like text (Board Titles, LLM captions) and image embeddings through a Multi-Layer Perceptron (MLP).

While PinSage/ItemSage capture "social curation" (how users save items to boards), the Unified Encoder integrates that knowledge with "content understanding" (what the item actually looks like and describes)

“In cases where certain features are defined for one entity but not the other, we substitute them with zero, ensuring a consistent data input.”

These features are being fed as tokens into a transformer, right? In that case, why do they use 0 as a substitute as opposed to a <NULL> token or something similar?

Actually, textual features (processed via a Hash Embedder) and continuous features (pre-computed embeddings like PinSage and ItemSage) are concatenated and fed into a 3-layer MLP (Multi-Layer Perceptron).

Substituting with a vector of zeros is the standard mathematical way to represent "no signal" in a neural network layer.

Why do they use 3 different tokenizers? I’ve never seen this before so I’m curious why it’s not more widely used.

Most modern models use BPE (Byte-Pair Encoding) or WordPiece. These are "middle-ground" solutions that handle both whole words and fragments automatically, usually making explicit unigram/bigram/character-level tokenizers redundant in a single model.

Using a 1-million-word bigram vocabulary is memory-intensive. LLMs typically prefer a fixed, smaller vocabulary (e.g., 32k to 128k tokens) that can represent any word. Pinterest manages this bloat using a Hash Embedder.

I didn’t understand anything from the hashing paragraph. What’s going on there?

Hashing in embeddings is used to significantly reduce memory consumption, handle massive or dynamic vocabularies, and manage unseen categories without needing a pre-defined dictionary.

2-hash reduces the risk of collisions.

Learning weights for the interpolation allows variation in embeddings.

“The sum of all token embeddings and the embedding features are concatenated”

What embedding features are they talking about? And do they mean they’re adding the embeddings to each other or they’re adding the embeddings to the features somehow?

The "embedding features" refer to features like the Image Encoder output, PinSage and ItemSage.

They’re summing the embeddings of the different tokens (as described in the tokenizer). This aggregated embedding is concatenated with the features.

“the problem of learning query

and entity embeddings is treated as an extreme classification problem”

What do they mean by extreme classification problem?

Extreme classification problem means the model is being trained to predict a specific "correct" item out of a massive list of billions of possibilities.

The classes are the items.

What’s the difference between sampled softmax loss vs regular softmax loss? Why do they use the former vs the latter and what are the tradeoffs?

In a standard classification task, the model calculates a probability score for every single possible class in the catalog. The denominator requires a summation over the entire catalog C.

In Sampled Softmax, instead of summing over the entire catalog, the model sums over a small subset C' which typically includes in-batch positives and random samples.

Sampled Softmax is much faster and requires less memory; mathematically speaking, it is an unbiased estimator of the full softmax gradient. {Need more work to understand this better}

What’s the purpose of the logQ correction in the softmax loss?

When you use Sampled Softmax, you only compare a query against a small group of "negative" items instead of the entire catalog of billions. If you sample negatives based on their popularity (which is common to make the model "work harder" on difficult examples), popular items will appear as "negatives" in your training data far more often than rare items.

Without correction, the model would see a popular Pin like "Standard Kitchen Layout" as a negative thousands of times. It would eventually learn that this Pin is "bad" or should have a low score simply because it was used as a negative so frequently, not because it was actually irrelevant to the queries.

The method subtracts the log of the sampling probability (log Q(y|x_i)) from the calculated similarity score (the logit) before the softmax is applied. By subtracting this value, the model "compensates" for how likely an item was to be picked as a negative.

What is going on in the equations they provide after “This is crucial to ensure that popular entities aren’t disproportionately penalized”?

{Not prioritizing this for now}

“To increase training efficiency, we share the pairs in the batch across all tasks with the same dataset.”

What does this mean?

For tasks that are using the same entity pairs, we only load that pair into the GPU memory once. This way, the query encoder (DistilBERT) only has to run once for that query to generate an embedding used by both tasks.

“After testing various Cache Time-To-Live (TTL) periods”

What factors should be considered when determining the ideal TTL?

  1. Cache Hit Rate
  2. Storage Cost

In Table 1: Summary of the different training datasets, it seems like the different datasets are all of very different sizes - if we’re training the embeddings together, wouldn’t the entity pairings with larger datasets drown out the other pairings?

I think this is a reasonable concern. It’s not clear how they address it in the paper but oversampling could perhaps be used.

“A common challenge in recommendation systems is the popularity bias, where certain pins are overrepresented due to their high appeal. To counteract this bias, we impose a limit on the number of times the same pin can be paired.”

Let’s say they hadn’t done this - what specific negative consequences would occur?

In machine learning, the model updates its weights based on the sum of losses across a batch. If a single popular Pin appears in 50,000 training pairs while a niche Pin appears in only 2, the model’s gradients will be overwhelmingly influenced by the viral image. The model becomes an expert at representing "head" (popular) content but fails to learn the unique features of "tail" (niche) content. This directly contradicts Pinterest's goal of being a discovery engine for a "wide range of interests".

Discovery Collapse: users would be trapped in a feedback loop where they are only shown what others have already seen, making it impossible to uncover new, relevant content from the billions of items available.

I’m worried that using in-batch negatives for query-query pairs could lead to a lot of false negatives. With Pins, users can always open one Pin, close it and then open another Pin - if they don’t end up opening other Pins, it’s fair to assume those are negative.

I don’t think this is the case for suggested queries as those are likely to lead to rabbit holes, with users not returning to the original query to review the remaining suggested queries. Is this a reasonable concern?

I think this is a fair concern but they still saw a 44% gain, suggesting that the "signal" of the actual clicks was strong enough to overcome the "noise" of the ignored suggestions.

“We ensure no data leakage—possible due to the inclusion of engagement features such as engaged queries—by maintaining a 15-day separation between the end of the training dataset and the beginning of the evaluation phase”

I don’t get how data leakage applies here

Let’s say you want the model to predict if a user searching for "emerald green sofa" will click on Pin A. If Pin A’s feature list already includes "emerald green sofa" because it was a historically engaged query, the model doesn't have to learn what "emerald" or "green" looks like. It just performs a simple string match. The model looks perfect during training but fails in the real world when a new, slightly different query (e.g., "forest green couch") appears, because it never actually learned the visual semantics—it just learned to read its own notes.

“This metric denotes the likelihood of the occurrence of the engaged entity within the top 10 entities when these entities are sorted in

descending order based on their similarity to the query.”

But this ignores the reranking stages, right? My understanding is OmniSearchSage is being used as part of a pipeline so the top 10 entities according to OmniSearchSage scores aren’t the Top 10 displayed to users; is that correct?

Yes. However, it’s still useful to consider this metric. If a Pin doesn't make it into the top group during the retrieval stage, the L2 reranker never sees it. Therefore, high Recall@10 at the retrieval stage is a prerequisite for a high-quality final experience, even if the final UI looks different.

During model development, engineers need to know if the encoder is improving. If they only measured the final UI "Top 10," they wouldn't know if a failure was caused by OmniSearchSage (Retrieval) or the subsequent L2 model (Ranking).

“Table 4 illustrates the impact of adding different text enrichments on the model’s performance.”

Don’t ablation studies depend on the order in which features are considered?

Yeah, but I guess the point of this table isn’t to isolate each of the components’ exact impacts but to show that they’re all useful.

“This aligns with general notions about multi-task learning: low-data tasks are unlikely to see regressions from multi-task learning”

Why don’t low-data tasks usually see regressions from multi-task learning?

Low-data tasks benefit from Multi-Task Learning (MTL) because they "eavesdrop" on the vast information processed for larger tasks.

Small datasets lack the diversity to teach a model complex features (e.g., the visual concept of "modernism" or "texture"). The 1.5 billion Pin pairs teach the encoder universal representations of images and language. The low-data tasks (like Product search) simply "use" these high-quality features that they could never have learned on their own.

“We train a model that comprises only the query and unified pin and product encoder. Subsequently, this model is compared with another model that fully incorporates all the encoders. Interestingly, there is almost no noticeable degradation in the metrics of the learned encoder, thereby essentially achieving seamless compatibility of the query embedding with pre-existing embeddings at no substantial cost.”

What is the difference between the two models being compared here?

The second model is trained to be "backward compatible." It must match the new Query Encoder results to the Unified Encoder and to the legacy PinSage/ItemSage embeddings simultaneously.

Model 1 (Simplified): Contains only two active pieces—the Query Encoder and the Unified Entity Encoder

Model 2 (Comprehensive): Contains the Query Encoder and the Unified Entity Encoder, plus the two Compatibility Encoders (the projection layers) shown in Figure 1. These extra layers are what "translate" the new query vector so it can talk to the old legacy systems.

“Figure 4: A simplified depiction of the search retrieval and ranking stack at Pinterest highlighting the integration points for OmniSearchSage embeddings.”

I don’t really understand what’s going on in this figure. I get that search will consist of retrieval, ranking (L1) and reranking (L2) but how does that connect to all the different stuff in the figure?

HNSW is an ANN algorithm used to retrieve the most similar items to the query.

The Embedding Index and the Inverted Token Index represent semantic and lexical similarity respectively.

“For this evaluation, we selected a set of 300 queries, deliberately stratified across both head and tail queries.”

Why do they stratify across head and tail queries?

Head Queries: These test the model's ability to memorize and retrieve high-signal, historical engagement. Success here depends on how well the model stores known associations.

Tail Queries: These test the model's ability to generalize. Since there is little to no historical data for a niche query (e.g., "blue velvet victorian cat throne"), the model must rely on its deep semantic understanding of the words and images to find a match.

Pinterest’s core value proposition is "discovery." Users often go to the "tail" to find unique, personalized inspiration. By measuring the tail separately, they ensure the OmniSearchSage encoder is fulfilling the company's primary mission of helping users find things they didn't know they were looking for.

Why are they using token-based retrieval as opposed to SearchSage in the comparison with OmniSearchSage?

That comparison seems primarily geared towards illustrating the utility of semantic search vs lexical search.

What are the 3 different categories they’re measuring in “Table 7: Online A/B experiment results of OmniSearchSage in Organic Search.”?

Retrieval, ranking and re-ranking. This table shows that the major gains from OmniSearchSage came in the retrieval and re-ranking stages.

Why do the impacts in Table 7 and Table 8 seem so much smaller than the improvements they talked about earlier in the paper like Table 2?

Offline: OmniSearchSage is tested in a vacuum. The metrics only reflect the quality of the encoder's vectors.

Online: OmniSearchSage is just one of hundreds of signals. Even if the encoder is 44% "smarter," it still has to pass through the L1 Ranker, the L2 Reranker, Whole Page Optimization, and UI constraints. By the time the results hit the user’s screen, the "purity" of the embedding’s impact has been diluted by dozens of other filters and business rules.

What is “Product Ads Retrieval” in Table 8 and how is it different from the other two rows?

This refers to the initial stage of the advertising pipeline where the model's query embeddings are used to find relevant product-based advertisements (Ads) from a massive catalog of ad candidates.

“One example of this at Pinterest is

interest classification, where we classify queries into a hierarchical taxonomy”

What’s the purpose of this classification?

These can be used as features for downstream ranking.

How would I convince a business leader that this project was worth doing?

Increased ads conversion means more money directly.

Increased search relevance and engagement means a better user experience that leads to improved retention and level of activity, plus word-of-mouth.

How would you improve this system if you were building it today?

  1. Replace the multilingual DistilBERT query encoder with a specialized 1B–3B parameter LLM backbone (like a pruned Llama-3 or Phi-4 variant).
  2. Replace the 3-layer MLP entity encoder with a Cross-Attention Transformer layer.
  3. Move from LLM-generated synthetic captions to a Native Vision-Language Model (VLM) like SigLIP-2
  4. Hard Negative Mining: Instead of just random in-batch negatives, I would use an LLM-driven negative sampler.
  5. Replace the 30-day static cache with a Dynamic Vector Refiner. Use a lightweight "Residual Layer" on top of the cached query vector that adjusts the embedding based on the user's last 5 clicks. This provides real-time personalization without needing a full GPU re-inference of the entire query
  6. Use a high-end LLM (like Gemini 1.5 Pro or GPT-4o) as an "Agentic Judge." It reviews the top 10 results and scores them on Visual Harmony and Logical Cohesion. This catches the "uncanny valley" results that math metrics often miss.

Level of understanding: 😐

2025 Paper original

How we built ATLAS: a benchmark for spend classification in scope 3 carbon accounting

LLMs are a popular approach to spend classification. However, no benchmark dataset of spend classification exists that allows systematic inter-comparison of different modeling approaches.

My notes

Blogpost

Summary

Context

Watershed wants to help companies measure and reduce their emissions. Scope 3 emissions are all indirect greenhouse gases produced in a company’s value chain—both upstream and downstream—not included in Scope 1 (direct) or Scope 2 (purchased energy). Scope 3 emissions account for the majority of most companies' carbon footprints.

Among companies reporting their value chain emissions, 70% rely on "spend-based" methods to calculate these impacts; spend-based methods involve multiplying financial expenditures on goods or services by industry-average emission factors. This calculation thus requires spend classification (mapping millions of financial transactions to specific economic sectors).

Problem

LLMs are a popular approach to spend classification. However, no benchmark dataset of spend classification exists that allows systematic inter-comparison of different modeling approaches.

Solution

ATLAS, the first spend classification benchmark dataset, along with the performance metrics of four baseline models.

Some key insights/new things I learned

  1. Degree of Fabrication: describes exactly where an economic activity sits in the value chain, ranging from raw materials (low DoF) to finished products (high DoF). The authors map all 295 BEA codes into simplified DoF categories, such as Raw Material, Semi-finished good, Finished good, Services, and Utilities.

What I didn’t understand or am unsure about

Question

Answer

What do they mean by high quality emission factors?

Accurate, representative, and scientifically validated coefficients used to convert activity data (such as fuel consumed or miles traveled) into greenhouse gas (GHG) emissions

“We generate the synthetic mappings based on metadata about each BEA sector A.3 using Claude 3.5 Sonnet to match the label distribution of an internal private dataset of de-identified company spend accounts”

What does this mean? Is it just that they initialize a dataset with a label distribution matching the private dataset and ask the LLM to generate synthetic account names that could map to those labels?

I think this is pretty much what they mean although the initial generated names are modified in a subsequent step.

“We then modify the account names with rules observed in the private dataset.”

Why do they modify the account names for and according to what rules specifically?

The goal is to make the account names more reflective of the private dataset than the initial generations might be. Real data is probably a lot messier and may use certain abbreviations or acronyms that the initial generation step might not have fully captured.

Unclear what rules they could use to address this though. I would’ve thought it’d be easier to include these rules in the synthetic generation phase as opposed to applying them afterwards.

“The DoF helps understand the approximate position in the value chain and highlights larger errors in reasoning.”

How does the DoF help highlight reasoning errors?

I’m not entirely sure but I think it might be because the LLM generating a DoF mismatch implies a misunderstanding of the expenditure's nature.

“Expenses related to finished goods and services are relatively dominant, whereas, purchases of utilities and logistics seem under-represented.”

They seem under-represented relative to other categories in the dataset or relative to their share in other datasets?

I think they mean relative to other categories in the dataset.

For the text embedding approach they propose, could they have benefited from a two-stage approach (bi-encoders for ranking + filtering and cross-encoder for ranking)?

Yeah, probably. Not sure if it would’ve done as well as the LLM-based approach but I think it’s worth evaluating.

It seems a little surprising that logistic regression performed so well, especially on the private benchmark. Why was it so effective?

Unlike the text-embedding or zero-shot LLM models, the logistic regression model was explicitly trained on the training set.

If the rules from the private dataset use a consistent vocabulary, a frequency-based model will naturally perform well by "memorizing" those patterns during training

What is meta-prompting?

Treats the AI as a prompt engineer to create structured, high-performance instructions, essentially using AI to determine the best way to ask for a result.

“We dynamically retrieve 50 examples from the training dataset which are incorporated into the prompt.”

How do they do dynamic retrieval? Pulling examples from the training set where the account embeddings were similar to the current account embedding being considered?

Yes, maybe using the same embeddings they mention in the text embedding approach.

Could an ensemble approach have helped? Perhaps hybrid rank fusion from the Top 3 of logistic regression and LLM with few-shot?

Yes, I think so since it seems likely that they have complementary strengths (fail in different ways) since their architecture is so different.

How would I convince a business leader that this project was worth doing?

ATLAS allows people outside the company to attempt to solve the same problem Watershed is trying to solve. If external researchers are able to improve performance on ATLAS, those improvements can be used to improve Watershed’s internal approach.

How would you improve this system if you were building it today?

  1. Beyond account names, include vendor-level metadata (e.g., website descriptions or SIC codes)
  2. Incorporate transaction frequency and amount
  3. Hypothetical Document Embeddings (HyDE): Since ledger entries are "terse", have an LLM "expand" the entry into a detailed description before embedding it
  4. Two-Stage Retrieval (Bi-Encoder + Cross-Encoder)
  5. Agentic Verification Loop: second LLM agent reviews the top-1 prediction and checks it against the Degree of Fabrication (DoF) to catch "larger errors in reasoning" before final output
  6. Use a loss function or selection logic that prioritizes reducing the Emissions Factor (EF) RMSE. This ensures that when the model is wrong, the carbon accounting error is minimized
  7. Conformal Prediction to provide a "confidence score." If the score is low, the system automatically routes the transaction to a human analyst

Level of understanding: 🤓

2026 Paper original

Agentic Multi-Source Grounding for Enhanced Query Intent Understanding: A DoorDash Case Study

Intent ambiguity in queries makes providing useful results challenging; traditional classifiers force a winner-takes-all assignment, while general-purpose LLMs hallucinate unavailable inventory.

My notes

Summary

Context

When users enter a search query on DoorDash, they could be looking for a restaurant name, a specific dish, a grocery item, a generic retail item.

Problem

Intent ambiguity in queries makes providing useful results challenging; traditional classifiers force a winner-takes-all assignment, while general-purpose LLMs hallucinate unavailable inventory.

Solution

Agentic Multi-Source Grounded system that addresses both failure modes by grounding LLM inference in (i) a staged catalog entity retrieval pipeline and (ii) an agentic web-search tool invoked autonomously for cold-start queries. Rather than predicting a single label, the model emits an ordered multi-intent set, resolved by a configurable disambiguation layer that applies deterministic business policies and is designed for extensibility to personalization signals

Design Details

  • Data
    • N/A
  • Architecture
    • Ask LLM to output either 1 or 2 intents based on
      • Items retrieved from semantic and lexical search for the user query in the catalog
      • Agentic web-search tool to retrieve real-time world knowledge
    • Disambiguation layer will narrow it down to 1 intent
      • Their paper uses historical win rates between the intents, if both the intents are in the conflict pairs data
  • Evaluation
    • Metric: Accuracy
    • Data: four benchmarks representing distinct traffic slices (𝑁 ≈ 30,000): Branded* (𝑁=7,809): retail brand-name queries; Retail (𝑁=7,988): general non-food retail queries; Tail (Synthetic) (𝑁=2,335): rare long-tail queries; and SOT (𝑁=12,651): a representative sample of global traffic across Head, Torso, and Tail segments
  • Deployment
    • {Batch and cache}
  • Post-deployment
    • N/A

Some key insights/new things I learned

  1. Misrouting at the intent determination stage will wreck your performance, regardless of ranking quality. If a restaurant shares the same name as a grocery item, guessing what a user means correctly is the primary determinant of search relevance.
  2. Token-set overlap (handling reordered terms) and partial matching (handling substrings)

What I didn’t understand or am unsure about

Question

Answer

“Traditional supervised multi-class classifiers [17] force a single-label assignment. While such models can emit top-𝑘 labels at inference, the softmax training objective encourages competition for probability mass, suppressing concurrently plausible intents under ambiguity”

What does this mean? Isn’t encouraging competition for probability mass what we want here?

If we’re more sure about one specific category, shouldn’t we reduce the number of results from other categories?

The paper doesn’t explain this super well but my best guess is that their approach works better because

  1. Even though it similarly returns one intent at the end, it retrieves catalog items from potentially any of the relevant categories to influence help it decide and performs agentic web search to gather more info if needed.

In other words, competition for probability mass is inevitable because we’re only returning one intent. However, their approach allows us to postpone evaluating this competition until we have access to relevant evidence (and whatever else we want to include in the disambiguation layer).

In equation (1) in Section 2.1 where they say “seeks the single category 𝑣∗ that maximizes the conditional probability”,

What is P(v|q)? Is that just the probability that the user’s intended category was v given the query was q?

The way they phrase it honestly seems unnecessarily convoluted.

Yes, that’s what it is. It’s convoluted but I guess in papers, you have to very pedantic about how you define these things.

“We optimize for the smallest set 𝑆 ⊆ 𝑉 such that the probability of the user’s true

intent 𝑣∗ being contained in 𝑆 is maximized”

It seems like there are two competing objectives here - minimizing the size of S and maximizing the probability that the user’s true intent is contained in S.

How do they handle this? Is it just by Satisficing and Optimizing?

Yes. They probably picked 2 because they assume that there are at most two primary plausible interpretations for the vast majority of ambiguous queries.

By including the multi-source evidence E in the equation, are they just saying that we should find the best S given the query and other evidence we have?

Yes.

“Query 𝑞 and catalog entities 𝑒 ∈ E are mapped to a shared embedding space via an encoder model”

What model do they use/should they use?

Not specified in the paper but they can use

  • Base Architecture
    • Two-Tower (or Bi-Encoder) architecture
    • Examples: BGE-M3, GTE-Qwen or SBERT
  • Fine-tune encoder using Contrastive Learning on click-through data to "pull" queries and correct items closer together in the embedding space

What is the prompt and what is the format for the catalog items inputted to the LLM? Just the text description or could we include images as well? Would the model need to be different for multimodal?

Unclear from the paper but it would probably be something along these lines:

  1. System Prompt
  2. User query: {user query}
  3. Evidence
    1. Retrieved Catalog Items:
      1. Item 1: {Entity Name, Category/Vertical}
      2. Item 2: …
    2. Web Search Results
      1. Summarized snippets from the autonomous web-search tool used for "cold-start" queries
  4. Policy Context

My guess is that including the images isn’t that helpful and probably isn’t worth the latency/cost increase required for a VLM like Gemini 1.5 Pro or GPT-4o. This shouldn’t be too difficult to actually test though.

Where does alpha come from in the fuzzy match equation?

It’s a configurable hyperparameter. They don’t specify how they determine the optimal value but they could conduct offline grid searches or A/B testing on historical logs to find the balance that maximizes precision without sacrificing too much recall.

What’s the difference between token-set overlap (handling reordered terms) and partial matching in their fuzzy match equation?

Token-set ignores word order, focusing on the set of words present then compares the overlap.

Partial matching looks for the best-matching segment. It takes the shorter string (usually the query) and slides it across the longer string (the entity) to find the maximum possible similarity for any substring

“we re-rank using a weighted fuzzy matching score”

Why use purely lexical reranking as opposed to Hybrid Rank Fusion? Or do they mean they’re using it just for filtering, not ranking?

They’re definitely using it for filtering but since they explicitly say they use it for reranking, they’re probably using both.

I don’t have a great explanation for why use purely lexical reranking as opposed to some sort of hybrid approach.

Since it’s only a small part of the pipeline though (retrieving evidence to help the LLM disambiguate intent), I think it doesn’t matter too much.

“For cold-start queries absent from the catalog, the model autonomously invokes an agentic web-search tool to retrieve real-time world knowledge.”

What is the agent actually looking for when it uses web-search?

It can search for something like "What is {user query}?" or "{user query} business type."

Then, it can look at the top search results and then summarize/format them to provide to the LLM.

What do they mean by “Strategic business rules” in 2.2.2 Agentic External Search and how do they “[enable] rapid policy updates without architectural changes?

Something like adding “Prioritize Retail for seasonal queries” to the prompt.

You don’t need to change the pipeline for this, just the prompt.

What are potential examples of “strategic overrides” in the policy context?

I think this is the same thing as the previous question but not completely sure.

I don’t understand the Pairwise Override function used when the model predicts a dual intent.

When they say “whitelist of conflict pairs for which historical data indicates the secondary intent should take priority”, do they just mean pairs where the secondary intent has more queries associated with it historically than the primary intent?

If yes, where do they get the labels for historical data from?

Look at historical data where the user was shown options from both intents and calculate in what % of cases an item from each intent was selected.

One ambiguity is that this would depend on users actually seeing multiple intents in the historical data. And if their previous architectures, similar to the one they use in this paper, always only showed one intent, there wouldn't be any historical data to use.

Is returning results from only one intent actually optimal? I think this may already have been improved on by others in industry - https://research.atspotify.com/2025/9/you-say-search-i-say-recs-a-scalable-agentic-approach-to-query-understanding

I don’t know why they do it this way; maybe blending the results is too confusing/complicated from a UI perspective.

I think it's worth at least experimenting with a multi-intent system, especially if they said earlier that correct intent classification is the biggest determinant of a successful search.

How could the personalization signals or learned re-ranker be incorporated into the disambiguation layer? What would this architecture look like?

We can collect signals like

  1. User History (Long-term): "User orders Grocery 80% of the time on Sundays."
  2. Contextual (Short-term): Time of day (Breakfast vs. Late Night), User Location (Home vs. Work).

Then just use XGBoost or a multi-task ranking neural network like MMoE or Progressive Layered Extraction.

Their evaluation is only focused on improvements in the intent classification layer, right? If that’s the case, could improvements in intent classification have come at the cost of worse overall results?

Yes, it’s worth looking at A/B results for metrics like Search Conversion rates.

In the abstract, they say “The system is deployed in production, serving over 95% of daily search impressions” but later they say “[The current architecture’s] offline batch-and-cache design limits

coverage to previously seen queries.

This means that 95% of queries are batch-and-cached queries, right? Doesn’t that seem unusually high? Do they use some sort of fuzzy caching?

I think it’s probably reasonable; you can batch a very large amount of stuff offline. Plus, query normalization will help.

Also, every time you encounter a new query, you add it to the cache. Long-term, the volume of never-before-seen queries will probably be pretty low.

Their analysis of their ablation study implies that Catalog Grounding was the biggest driver of improvements, followed by Agentic Search and Dual Intent Disambiguation. However, couldn’t these findings have been different if they’d done the ablation in a different order?

Yeah, unless I’m missing something, this seems a bit misleading.

Where did the labels for the evaluation data come from (besides the synthetic data)? Manual labeling or inferred intent from search resolution?

Unclear but I think either approach could work.

How would I convince a business leader that this project was worth doing?

  1. Provide specific examples of real-life searches that were resolved correctly with the new system whereas the old system failed.
  2. Conversion rate connected to $ impact
    1. For any given user search, either what they’re looking for exists or doesn’t exist on the platform. If it exists, either we show it to them or we don’t. If it exists but we don’t show it to them, some percentage of users won’t try to reword their query. These are potential purchases that are lost and users whose experience with DoorDash has been suboptimal, diminishing their interest in future purchases (this is also true for users who end up rewording their query).
    2. Lost Revenue = Total Searches * Intent Failure * Bounce Rate * Avg. Order Value

How would you improve this system if you were building it today?

  1. Blended results from multiple inferred intents (similar to Spotify paper)
  2. Knowledge Graph augmented RAG
    1. When a user searches for a specific brand, the system traverses the graph to find the exact brand family.
  3. Distilled SLM fallback for cache misses

Level of understanding: 🤓

2024 Blogpost original

From grep to SPLADE: A Journey Through Semantic Search

Semantic search is not ideal for literature search because it has limited interpretability and it is potentially non-deterministic, especially at a larger scale.

My notes

Summary

Context

Elicit is a tool for systematic literature review. The first step in that review is a search that looks similar to what you would see in Google Scholar.

New! Automate your Systematic Review with Elicit

Problem

Semantic search is not ideal for literature search because it has limited interpretability and it is potentially non-deterministic, especially at a larger scale.

Solution

Query expansion using SPLADE

Some key insights/new things I learned

  1. SPLADE models can be fine-tuned on the domain
  2. Automated comprehensive search

What I didn’t understand or am unsure about

Question

Answer

How does OpenAI’s embedding API work? What model/architecture does it use?

The two main OpenAI text embedding models are text-embedding-3-large and text-embedding-3-small.

Key Architectural Components

  1. Transformer Encoder: the core architecture is similar to BERT or RoBERTa
  2. Matryoshka Representation Learning (MRL): earlier dimensions hold the most critical information, while later dimensions add refinement. This allows practitioners to reduce embedding dimensions (e.g., from 2048 to 128) for efficiency, with minimal loss in accuracy
  3. Pooling Layer: aggregates individual token-level embeddings produced by the transformer into a single vector representing the entire text chunk.
  4. Training Process; likely via contrastive learning techniques

Are there really no good ways to make standard semantic search interpretable for a regular user?

There are a few decent ones:

  1. Attribution Highlighting: use Cross-Encoders to find the specific passage that most closely aligns with the query's intent and highlight it
  2. LLM-Generated Justifications: use a fast LLM to "read" the top results and provide a 1-sentence explanation

How does SPLADE actually work?

How SPLADE works:

  1. The input is passed through a transformer (like BERT). For each token i in the input, the model produces a contextualized hidden representation h_i taking into account the full sentence.
  2. SPLADE projects each hidden state h_i onto the entire vocabulary (usually ~30,000 terms) V using the pre-trained Masked Language Model (MLM) weights. For for every possible word j in the vocabulary, we calculate

w_ij = transform(h_i) E_j + b_j

  1. E_j is the context-independent embedding of token j from the initial embedding layer of the model.
  2. When we calculate transform(h_i) E_j, we are essentially measuring the cosine similarity (unnormalized) between the contextualized meaning of the current word and the general concept of every other word in the dictionary.
  3. The transform layer is there to translate h_i into the same space as E_j. The actual equation is


transform(h_i) = LayerNorm(GELU(h_i * W_dense + b_dense)

  1. {More work needed to understand this math}
  1. Once we have w_ij for every (input token, vocab word) pairing, these individual distributions are combined using Max Pooling across all tokens. SPLADE takes the maximum score for each vocabulary word across all input tokens.


w_j = max w_ij for i in [1,n]

  1. Log-Saturation


Final_weight_j = log(1 + RELU(w_j)

  1. Pre-trained SPLADE models are already trained with the FLOPs loss which adds a penalty to the total loss function that "punishes" the model for having too many active (non-zero) dimensions.
  2. The final output is a sparse vector with weights. The weights generated by SPLADE are used to calculate a relevance score between a query and a document via the dot product of their respective sparse vectors.
    1. Because the vectors are sparse, the system does not calculate the score for every document in the database. It uses an inverted index to find only the documents that share at least one non-zero term with the query.

How is using SPLADE better than prompting a generic LLM to provide you all variations for search terms?

SPLADE outputs a sparse vector that looks exactly like a traditional search index entry. With prompting, the output is raw text. You then have to re-tokenize, re-clean, and re-structure that text into a format your search engine understands.

How does SPLADE know which words/phrases to expand? Their graph shows some sort of entity recognition but how does it actually do that?

That’s just a visual abstraction, not how SPLADE actually works.

As explained in the previous answer, for every token in your query, SPLADE generates a weight for the entire vocabulary. It then uses a max-pooling operation. If "cardiovascular" is highly related to both "heart" and "health," it takes the highest score for that word across the whole query.

Is there a good way to combine BM25 with SPLADE or should we only use one?

You can do either of the following

  1. Use SPLADE to generate the terms, and let BM25 handle the ranking of those terms (updating the BM formula to include the SPLADE weights).
  2. Parallel Hybrid Retrieval with Reciprocal Rank Fusion or Weighted Convex Combination
    1. If we have evaluation data, we can experiment with the different approaches

How exactly would training a custom SPLADE model on a particular domain be done?

  1. Domain-Adaptive Pre-training (DAPT): Take a base BERT-style model and run it through Masked Language Modeling (MLM) on a large corpus of raw domain text (e.g., 1 million legal briefs). This "teaches" the model your jargon.
  2. Generate Triplets: Create a training set of (Query, Positive Document, Negative Document). If you lack labeled data, use an LLM (like GPT-4) to generate synthetic queries based on your documents.
  3. Select a Teacher (Distillation): Use a powerful, slow Cross-Encoder to score your document pairs. You will train the "smaller" SPLADE model to mimic these scores.
  4. Fine-Tune with Sparse Losses:
  • Ranking Loss: Use SparseMarginMSELoss to align with the teacher.
  • FLOPS Regularization: Apply a mathematical penalty (the "sparsity dial") to force the model to use as few tokens as possible.
  1. Hard Negative Mining (ANCE): Use your current model version to find "convincing" wrong answers (hard negatives). Retrain the model on these specifically to improve its ability to distinguish between very similar technical concepts.
  2. Evaluate for Sparsity: Unlike standard AI, you must validate two metrics: nDCG (accuracy) and Sparsity Ratio (ensuring the vectors aren't getting too "heavy" or unreadable for the user).

Is semantic search’s non-determinism because of the use of ANN or are there other reasons as well?

ANN Index Construction is the main reason but other reasons include

  1. GPU Floating-Point Non-Associativity
  2. Atomic Operations & Race Conditions
  3. Inference Precision

What are some good metrics for evaluating SPLADE vs other approaches (both general and those especially relevant for the specific context of literature review)?

Offline

  1. Determinism: whether the same query produces the identical result list every time
  2. nDCG (Ranking Quality): rank-weighted aggregation of relevance of top results, normalized by the true distribution of relevant items available for that specific search
  3. Recall@K

Online

  1. Task Success Rate (TSR): percentage of users who save/export a relevant paper within a 10-minute session

How would I convince a business leader that this project was worth doing?

  1. Recall@K is a relatively explainable metric but it might not actually improve if we use SPLADE
  2. Task Success Rate
  3. Search speed?
  4. Demo of searches conducted using SPLADE vs original methodology to demonstrate interpretability improvements
  5. Cost savings from less memory and storage?

How would you improve this system if you were building it today?

  1. On-Policy Distillation of large SPLADE model into smaller model
  2. Agentic search
    1. The agent could follow citation trails
    2. Iterative Query Refinement based on intermediately retrieved results
    3. Reasoning Logs for transparency

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Learn more about SPLADE and search indices
2024 Blogpost original

Improving Retrieval on Ramp with Transaction Embeddings

Given that the majority of employees are not finance and accounting experts, accounting coding can be confusing and error prone for them. Employees have to search through hundreds of different accounting codes, with little clarity on what…

My notes

Summary

Context

Accounting coding is the process of categorizing transactions based on each business' accounting system; employees of Ramp customers are responsible for coding their own transactions and expenses.

Problem

Given that the majority of employees are not finance and accounting experts, accounting coding can be confusing and error prone for them. Employees have to search through hundreds of different accounting codes, with little clarity on what they exactly represent. As a result, any human error during this process results in additional work for the finance teams when closing the books.

Solution

Generate transaction embeddings that could cluster similar transactions by their respective GL category. Then, predict a GL category for a new transaction by matching accounting codes attached to similar transactions.

Some key insights/new things I learned

  1. Representing transactions
    1. Transactions can consist of features like Merchant name, Merchant category name (MCC), Department name, Location name, Amount, Memo, Spend program name, Trip name
    2. Input format: represent transactions as strings using the features above. Example:
      “Name: Beirut Bakery, Category: Restaurants, Cafetaria, Department: Engineering ….”
    3. Learned representation: Starting with a pre-trained encoder as a base model, fine-tuned it on Ramp’s transactional data.
  2. Loss function (Triplet loss)
    1. Evaluates samples in triplets, consisting of an anchor, a positive sample (same GL code as anchor), and a negative sample (different GL code as anchor)
    2. Triplet loss learns to “push” the embeddings of negative samples away from the anchor, and pull positive samples towards the anchor at the same time.

What I didn’t understand or am unsure about

Question

Answer

Where did the labels come from?

The specific labels used for the training dataset seem to be the final accounting codes assigned to each transaction by the finance teams using the platform.

“We represented transactions as documents with attached labels (chart of account codes).”

By ‘document’, do they just mean a string?

Yes.

What are some examples of GL categories? Are GL categories hierarchical?

Examples; “7000 Personal Expenses”, “6300 Utility Expenses”, “Travel: Sales”

The structure of these GL categories seems quite varied; this might be because GL categories are company-specific and different companies have different ways to structure them.

In at least some cases, GL categories are indeed hierarchical; the post mentions “Travel: Sales” and “Travel: Engineering”, which are both subsets of “Travel”

“the model can … cluster transactions by their respective GL categories.”

What does it mean to cluster transactions by their respective GL categories?

My interpretation is that transactions with the same GL categories would have similar embeddings. Is that what they mean?

Yes. In this context, "clustering" does not necessarily refer to a secondary algorithm (like K-Means). Instead, it refers to the spatial organization of the embedding space. Because the model is trained to minimize the distance between transactions with the same label, those transactions naturally group together into distinct "neighborhoods" or clusters.

“Though there are many off-the-shelf embedding models and various approaches to training custom embedding models, we chose an architecture that could easily scale with the large variety of transactions on Ramp.”

What specific criteria besides embedding dimensionality should be considered in a model’s architecture to fulfil this business need?

  1. Encoder-based foundation: efficient at turning string-based "contextual transactions" into dense vectors.
  2. Low Inference Latency: since suggestions are surfaced in real-time dropdowns while a user is typing or viewing a transaction, the model must be lightweight enough to generate an embedding in milliseconds.
  3. Computational Cost: a smaller, encoder-only architecture (like BERT-base or DistilBERT) is much cheaper to run at the scale of "millions of transactions per month" compared to larger LLMs.
  4. Compatibility with ANN Search: the architecture must produce embeddings that work well with Approximate Nearest Neighbor (ANN) indexing, allowing for sub-second searches even as the dataset grows.

Standard pre-trained encoders (like BERT) usually have fixed dimensions. If they needed something smaller, what could they do?

  1. Choose a smaller base model: use a variant like MiniLM or TinyBERT which natively has a smaller dimensionality.
  2. Add a projection layer: add a custom linear layer at the end of a larger model to compress the output down to their desired size (e.g., 256) during the fine-tuning process.

“We then use cosine similarity as a distance metric in the learnt latent space to quantify the similarity between two transactions.”

Is cosine similarity the only useful distance metric they could have used or were there other reasonable alternatives that might have been worth exploring?

Other options include Euclidean Distance, Dot Product and Manhattan Distance.

The first two are the equivalent to cosine similarity if the vectors are normalized (which they most likely are).

Manhattan Distance can make the system more robust to outliers but isn’t as popular in industry. {Need more research to understand why}

If codes are hierarchical, do negative samples include only those with 0 overlap? In other words, would “Travel:Sales” be considered a negative sample if the anchor is “Travel:Engineering”?

Yes. These sorts of pairs are useful as Hard Negatives to help the model learn fine-grained distinctions.

In theory, we could use Hierarchical Triplet Loss or Weighted Triplet Loss to treat partial matches differently from zero-overlap pairs. I guess they don’t do that because a partial match isn’t useful for them.

How exactly does triplet loss learn “to “push” negative samples away from the anchor, and pull positive samples towards the anchor”?

What do the model updates look like exactly and how does the margin fit into it? What effect does changing the margin have and how do you decide what margin to use?

Triplet Loss Function: L = max(d(A, P) - d(A, N) + alpha, 0)

  • d: cosine distance function = 1 - cosine similarity
  • A: anchor embedding
  • P: positive embedding (same label as A)
  • N: negative embedding (different label than A)
  • Alpha: the margin (a constant)

Gradient Updates

  1. Simultaneous weight updates
    1. The "Pull": To make the loss smaller, the model must minimize d(A, P). The gradient tells the model's weights to shift so that the coordinates of A and P move closer together in the vector space.
    2. The "Push": To make the loss smaller, the model must maximize d(A, N). The gradient tells the model to shift the coordinates of A and N further apart.
  2. The "Stop": If the negative is already far enough away (specifically, if d(A, N) > d(A, P) + alpha), the loss becomes 0. The model stops updating for that specific triplet because the goal has been met.
  3. Also stop when we’ve reached the set number of epochs/steps, regardless of whether we’ve reached the margin.

Margin

  1. Effect of different margin values
    1. High Margin: Forces the model to create very distinct, widely separated clusters. This is harder to learn and can lead to "model collapse" if the data is too similar to push that far apart.
    2. Low Margin: Easier to satisfy, but may result in clusters that are too close together, leading to "confusion" or misclassification when a new, slightly ambiguous transaction appears.
  2. How to decide?
    1. Empirical testing. Start with a small value (like 0.2) and monitor the "Validation Accuracy" or "Cluster Silhouette Score.

How does "BatchSemiHardTripletLoss" (mentioned in the post) work differently than standard triplet loss?

Standard triplet loss

  1. Pre-calculate triplets (Anchor, Positive, Negative) before training.
  2. Most random triplets are "Easy" (already satisfy the margin). The model spends 99% of its time doing math that results in 0 loss, learning nothing.

BatchSemiHardTripletLoss

  1. Instead of pre-picking triplets, the model looks at every possible combination within a single training batch on the fly.
  2. It uses the margin as a filter to pick the most productive examples.
    1. Easy (Ignored): N is already outside the margin. (Loss = 0).
    2. Hard (Ignored/Skipped): N is closer to the anchor than P. These are the "impossible" cases; the algorithm skips them to prevent the model from getting stuck on noise.
    3. Semi-Hard (The Focus): N is further than P, but still inside the margin.

“We found the BatchSemiHardTripletLoss function to produce best results.”

How do they measure the best results? Are there generally any tradeoffs to using BatchSemiHardTripletLoss?

They don’t explicitly say how they measure the best results but here are some options for metrics:

  • Suggestion Acceptance Rate: how often a user clicks the suggested code
  • Silhouette Scores: measure how tight clusters are
  • Recall@K (how often the correct category is in the top $K$ suggestions)

Tradeoffs:

  1. Pros
    1. Prevents "model collapse" where the model gets stuck on impossible/noisy examples (Hard Negatives)
    2. Every update is "meaningful" because it only looks at samples that almost break the rules.
    3. It is very efficient for large datasets.
    4. “Online" mining (doing it during training) is faster than "Offline" (pre-calculating triplets)
  2. Cons
    1. It may take longer to converge than "Hard" mining because the learning signal is less aggressive.
    2. You ignore "Easy" triplets, which means a large portion of your batch data might result in zero math being done.
    3. It is highly sensitive to batch size. If your batch is too small, you might not find any semi-hard negatives, and the model won't learn anything in that step.
    4. Finding these triplets on-the-fly requires calculating a "distance matrix" for every batch, which is computationally heavy for your GPU.

More details about how triplet loss functionally differs from contrastive loss, how this affects the generated embeddings and why that would lead to preferring triplet loss for this use-case.

Functional Differences

  1. Data Input: Contrastive loss works on pairs (either positive or negative). Triplet loss works on triplets (Anchor, Positive, and Negative simultaneously).
  2. Distance Logic: Contrastive loss uses an absolute margin (it tries to push negatives beyond a fixed distance). Triplet loss uses a relative margin (it only cares that the Negative is further away from the Anchor than the Positive is, by at least the margin amount).

Effect on Embeddings:

  1. Contrastive Loss: Forces all items with the same label to collapse into a very tight, singular point. It focuses heavily on separating "this" from "not this."
  2. Triplet Loss: Allows for intra-class variance. It lets "Travel: Sales" transactions stay somewhat spread out, as long as they are collectively closer to each other than they are to "Office Supplies."

Why Ramp preferred triplet loss

  1. Contrastive loss often hits local minima faster—it gets "stuck" once the classes are separated.
  2. Triplet loss "maintains more tolerance towards intra-class variance.", avoiding feature collapse (model over-optimizes on a few high-signal features, losing its ability to generalize).

Why feature collapse is an issue

  1. Original use-case
    1. Not necessarily a big issue but if we encounter a new data point which doesn’t contain the specific easy feature the model had optimized for, our predictions will suffer.
    2. If your clusters are collapsed into single points, every prediction looks equally "perfect" to the computer. Triplet loss lets us use distances to determine confidence levels for our prediction.
  2. Other use-cases
    1. Use-cases like analytics and suggested memos might require more nuanced embeddings {add explanation for why}.
    2. Anywhere we’re using RAG could benefit from more nuanced embeddings since we’d want to make sure what we’re retrieving is closely aligned with the query.

Do they use new lines vs spaces vs other options when formatting the string input to the sentence transformers?

Spaces are the standard delimiter for sentence transformers models, so it looks like "Title: [T] Description: [D]...”

Is it fine to have the amount and other numerical features included as strings in the input to the transformer? What are the alternatives if not?

I’m a little worried that numbers are so specific that we’ll end up with a lot of unique tokens and it’ll be hard for the model to learn their representations.

This seems pretty standard but I don’t have a great theoretical explanation for why this is fine; I guess the model can learn that a long string of numbers indicates a larger transaction than a smaller string.

Even when the number is split into multiple tokens, we still have positional encodings. Because of Self-Attention, a token's final representation is determined by its neighbors so the representation of a “1” followed by a “,” is different from a “1” that isn’t followed by a “,”.

We can round the transactions (logarithmic rounding might be best).

Alternatives:

  1. Late fusion: pass the text through the Transformer to get a text vector and separately, normalize the raw number for numerical features and concatenate with the text vector into a final hybrid embedding.

“Specifically, if we mine triplets at random, the model quickly learns the difference between the ‘Travel’ and ‘Repairs & Maintenance’ GL categories, and the loss function reaches a plateau.”

What does this mean?

Random sampling creates "Easy Negatives"—triplets that are so obviously different that the model learns nothing from them. When most of the samples are easy negatives, the amount of learning per epoch would be considerably lower.

“Using large batch sizes and the loss function allowed us to keep our training data relevant, making sure we were consistently feeding in informative triplet samples during training.”

What does this mean?

By “the loss function”, they’re referring to BatchSemiHardTripletLoss which, as we’ve already discussed, focuses on triplets that would be useful for learning. This “loss function” isn’t just a loss function but also a sampling filter.

Using larger batches is necessitated by this filtering to make sure we get at least one useful example in the batch.

“If the prediction confidence for a GL category is high, we can make a more overt recommendation to the user, like using that as the default value for a given transaction.“

How is the confidence threshold determined?

Determine accuracy requirement for overt decision and evaluate validation set to see how we can get there by conditioning on the following heuristics:

  • % of top-k neighbors with the same label
  • Average among top-k neighbors or closest distance
  • Difference between distance for top label vs 2nd best label

“Using these features and labels, we were able to … match querying transactions to similar, already-coded transactions to reference.”

How exactly do they pick the actual label for the new transaction? KNN?

They probably use KNN. They can calibrate K by evaluating the performance for different values.

What evaluation metrics could they have used to assess the performance of their overall approach?

  1. Retrieval/Ranking
    1. Recall@K
    2. Mean Reciprocal Rank (MRR)
    3. Hit Rate (Specifically for the "top-1" suggestion shown as the default value in the UI)
  2. Embedding Space Quality
    1. Silhouette Score: Measures how tight the clusters are and how well-separated they are from other categories.
    2. Davies-Bouldin Index: A ratio of within-cluster distances to between-cluster distances.
  3. Classification & Calibration
    1. Macro/Micro F1-Score
    2. Expected Calibration Error (ECE): Used to ensure that if the model says it has "90% confidence," it is actually correct 90% of the time.

Is there any risk of position bias or degenerate feedback loops?

Small risk; I think every transaction is being reviewed by finance teams which should limit the # of inaccurate labels used to train the model. However, finance teams are probably still somewhat biased towards accepting the user-assigned label.

They’d need to make sure the user-assigned labels aren’t going into the model training loop before finance teams have reviewed them.

How could transaction embeddings help with the following use-cases:

  1. Providing recommendations and insights on business spend and transactions
  2. Segmenting businesses by their spending patterns

Providing recommendations and insights on business spend and transactions

  1. Anomaly detection; transactions that land far from their expected category center are flagged as insights.

Segmenting businesses by their spending patterns

  1. Business Fingerprinting: A company can be represented as the centroid (average vector) of all its transactions. This "fingerprint" captures the unique DNA of how they spend money.
  2. Similarity Matching: Ramp can use these fingerprints to group similar companies (e.g., "Remote-first startups" vs. "Industrial manufacturers").
  3. Personalized Experience: Once segments are identified, Ramp can tailor the product experience, such as default spend limits or specific accounting workflows, based on what "companies like yours" typically use.

How would I convince a business leader that this project was worth doing?

Describe specifically how much time was actually saved on the accounting coding process; break it down by the different ways it was saved. These time-savings are valuable for Ramp customers, reducing churn, facilitating potential future upselling and creating positive word-of-mouth.

For the other use-cases, describing the impact we’d want to communicate requires more details.

How would you improve this system if you were building it today?

  1. Hybrid Tabular-Text Architecture (Late-Fusion Transformer)
    1. Pass text (Merchant/Memo) through a BERT-style encoder and numerical features (Amount/Date) through a continuous embedding layer (like a Periodic Activation network). Concatenate the resulting vectors before the Triplet Loss. This preserves exact numerical precision while keeping semantic context.
  2. Cross-Encoder Reranking for "Tricky Cases"
    1. Use the current embedding model to retrieve the top 20 candidates, then pass them through a Cross-Encoder.
    2. A Cross-Encoder processes the query and the candidate simultaneously, allowing for token-level comparison. This would significantly boost the accuracy of those "High Confidence" predictions they use for default values.
  3. Replacing Triplet Loss with InfoNCE (Contrastive Learning)
    1. InfoNCE treats every other sample in a large batch (e.g., 2048+ on modern H100s) as a negative simultaneously. This eliminates the need for manual "Semi-Hard" mining logic and provides a much smoother, higher-resolution gradient for the model to learn those subtle departmental differences.
  4. Entity-Resolution Pre-processing
    1. Use a small, high-throughput LLM (like Llama-3-8B) to perform Entity Resolution on the merchant string before embedding (e.g., converting "SQ *STARBUCKS #123" to "Starbucks").
    2. This removes "string noise" that consumes the model's attention, allowing the embeddings to focus purely on the behavioral patterns (Amount/Category/Location) rather than variations in merchant spelling.

Level of understanding: 🤓

2025 Paper original

You Say Search, I Say Recs: A Scalable Agentic Approach to Query Understanding and Exploratory Search at Spotify

Certain kinds of complicated user queries are not well-suited to the existing search architecture of Spotify.

My notes

Summary

Context

Users can search for content (songs, albums, artists, podcast, podcast episodes, audiobooks, playlists) on Spotify.

Problem

Certain kinds of complicated user queries are not well-suited to the existing search architecture of Spotify.

Solution

AI agents for query understanding and search.

Some key insights/new things I learned

  1. Pipeline
    1. Parallel Fusion Router
      1. Query is processed by an LLM router, which selects the most appropriate route based on the inferred intent. A route may involve either a single tool call or multiple parallel tool calls. Results of single calls are returned directly to users. Results of parallel tool calls are organized into separate SERP (Search Engine Results Page) sections.
      2. They use a pre-fusion configuration where multi-call routes are predefined as bundles.
    2. Tool-use
      1. Tools include traditional machine learning-based retrievers and rankers for search and recommendation, such as models that capture user-item or item-item similarities. Also include sub-agents (for example, technologies like AI DJ) that invoke LLMs to perform additional reasoning on the query or refine the results. These can be useful when fulfilling the intent requires external knowledge, complex reasoning, or a final SERP refinement step to enhance aspects like relevance or content diversity.
      2. Query rewriting for tool-use: for example, given a query “new indie rock releases,” the LLM router might enrich the request by adding temporal, genre, and entity-type filters, producing a structured call such as: {"route": "personalized_recs", "genre": "indie rock", "entity_types": ["track", "album"], "max_weeks_from_release": 4 }
    3. Post-training
      1. Teacher-student distillation with Rejection sampling Fine-Tuning (RFT)
  2. Evaluation
    1. Offline
      1. LLM-as-a judge approach to assess the quality of results across several dimensions, including relevance, diversity, and freshness.
      2. Test set consists of a blend of real and synthetic query-user pairs, designed to cover a wide range of use cases and routing strategies
    2. Online
      1. Rely on user interaction signals such as clicks and streams.

What I didn’t understand or am unsure about

Question

Answer

“Search systems have traditionally struggled with these types of exploratory searches, focusing on textual and semantic retrieval and on helping users find and play content they already know”

In what specific ways would an agentic search approach better help users find content they weren’t already familiar with vs SOTA standard approaches?

  1. With an agentic approach, queries like “new indie rock” could be restricted to results within a max_weeks_from_release threshold. SOTA standard approaches like hybrid search with personalized reranking are unlikely to be able to apply this sort of filter. They will surface either content with matching keywords (lexical similarity) or content whose meaning is similar to the phrase “new indie rock” (semantic similarity). This doesn’t properly capture the user’s intent to find new content and it will likely prioritize content similar to what the user has engaged with in the past.
  2. For a query like “Artists similar to Mozart”, lexical similarity would be almost entirely useless. Semantic similarity might help since artists similar to Mozart would probably have similar embeddings as Mozart but there’s a bit of noise because the search algorithm wouldn’t know to exclude Mozart himself and the embedding of the whole phrase is somewhat different from just “Mozart”. The personalization phase would then prioritize content similar to what the user has engaged with in the past, even though that’s not necessarily what they want. An agentic approach could route the query to item-item similarity tools and perhaps deprioritize personalization, helping the user to find new content.

“PFR strikes a balance between the flexibility of full orchestration and the efficiency of a simple router”

In what sense is full orchestration more flexible?

  1. Full orchestration can make multiple sequential calls, using the output of one tool to inform the decision for the next.
  2. Full orchestration can allow the LLM to dynamically select any combination of tools from its entire library at runtime.
  3. In a fully orchestrated system, the LLM can synthesize a completely original response by merging data from various tools into a unique narrative or summary. PFR is constrained by a Structured SERP (Search Engine Results Page), where every tool output is mapped to a specific, predefined visual section of the app's interface.

Let’s say in 99% of cases the best tool for the job was the original search approach that Spotify was using; in that case, would the additional latency from the LLM router coupled with a negligible benefit mean that the agentic approach is not worthwhile?

Most likely, but it would depend on the actual latency - high cache hit rates and really fast router LLMs would reduce latency considerably.

In any case, the evaluation performance shows that the benefits were fairly large.

The Search Engine Results Page could be of different sizes (and thus need different numbers of results needed to populate it) on different devices; how do/could they address this?

  1. Downstream tools could handle the specific "contextualization," such as adjusting the number of items returned based on the device's screen constraints.
  2. The LLM router could also explicitly extract and pass structural parameters (like result limits or filters) as part of the tool call formulation.

How is the SERP layout defined? Is this dynamic or pre-specified?

  1. The tool call routes are determined by the LLM based on the inferred intent.
  2. The routes are mapped to specific visual components of the SERP. These components are probably pre-defined by human engineers (albums, podcasts, etc).

“These different intents are fulfilled by triggering three parallel tool calls, with their outputs organized into separate SERP sections.”

How would you structure the tools and their docstrings to enable the LLM to accurately map between these intents and the appropriate tool calls?

  1. Tool Structure
    1. The tool's "Name" should be the Route ID (e.g., personalized_recs or artist_latest_content).
    2. The tool must include a Parameter Schema that allows the LLM to perform "facet extraction"—converting natural language into structured filters like genre, entity_types, and max_weeks_from_release.
    3. The parameters should be optional to allow the LLM to handle varying levels of query specificity (e.g., "new music" vs. "new indie rock from last week").
  2. Docstring Design
    1. The docstring should define the Exploratory Intent the route serves, such as "new releases for me" or "broad music searches"
    2. For multi-intent queries (like the Lady Gaga example), the docstring must explain that the route generates a fused response to address several plausible goals simultaneously (e.g., albums, singles, and release dates)
    3. To prevent mapping errors, docstrings should include "Boundary Definitions" (e.g., "Use this for broad artist queries; use navigational_search for exact song titles"

“PFR … improves accuracy for the most likely intent through better query understanding”

How is this evidenced by the text?

  1. It doesn’t seem like evaluation metrics improving was due to tool upgrades since they say the “tools include traditional machine learning-based retrievers and rankers for search and recommendation”.

“In the pre-fusion configuration, multi-call routes are predefined as bundles. This approach is faster and more efficient, as the LLM only needs to generate the parameters for a known route, minimizing the number of output tokens.”

I think I need more details about why this is actually faster.

  1. In a "post-fusion" or general agentic setup, the LLM must describe which tools it is using and how to combine them (often in a verbose JSON format). In pre-fusion, the LLM likely only outputs a Route ID (e.g., Route_1) and a few values. Since LLM latency is roughly linear to the number of tokens generated, cutting the "thinking out loud" part of the response directly slashes milliseconds off the response time.
  2. Because the task is simplified to picking a known route, Spotify can use a smaller, efficient LLM.
  3. The system prompt can also be shorter since the LLM is post-trained to know which tool bundles to use when.

“By passing only the query (without user features) to the LLM router, we maximize the cache hit rate. When needed, user features are forwarded to the downstream tools, which handle personalization and contextualization of the results.”

What are examples of queries where we can completely ignore user features?

“english lessons for spanish speakers podcast”, “top 50 songs in Japan”, “Tracklist for [Album]”

“They also include sub-agents (for example, technologies like AI DJ) that invoke LLMs to perform additional reasoning on the query or refine the results.”

In what circumstances would these subagents be triggered? How are these determined?

  1. When the user's intent requires information not found in the immediate search index (as determined by the LLM)
  2. If traditional retrievers return empty sets for a semantic query
  3. Confidence scores for intent which could be used to determine the need for sub-agents

“LLM router also determines the optimal parametrization of the tool calls”

What does this mean?

  1. The LLM inherently performs query expansion—adding related terms or constraints that the user didn't explicitly type but are necessary for a high-quality result. It is "optimal" because the LLM uses its reasoning capabilities to choose values that a standard keyword search might miss. For example, the LLM decides that "new" should specifically mean "released within the last 4 weeks" (max_weeks_from_release: 4) based on the context of the music industry, rather than a fixed, hard-coded rule.

“For instance, a broad, recommendation intent requires a page layout designed to support content exploration, so will involve richer visual cues and contextualization.”

What do they mean by “richer visual cues and contextualization”?

Figure 1 illustrates this by showing UI sections with album artwork, artist avatars, and distinct groupings rather than a single vertical list.

Contextualization refers to the information and metadata that explain why a result is relevant to the user's discovery journey. This likely includes "explanation strings" or headers that provide context for the discovery, such as "Because you like [Artist]," "Trending in [Genre]," or "Recently released." It tells the user the "reason" they are seeing an unfamiliar item, which is a key part of exploratory search.

“We address this by applying post-training techniques (Figure 2). The process begins with a training dataset composed of both real and synthetic examples.”

How would these real and synthetic examples be obtained respectively?

Real

  1. Sampled from Spotify's historical search logs. A "real example" at this stage is just a raw data point: (User X, Query: "vibey 80s"). It doesn't have a label for the routing decision yet, so it cannot be used to train the router directly.
  2. A powerful "Teacher LLM" is prompted to generate the correct routing decision for each example. To do this, it may take into account what action the user actually took after searching.
  3. The Teacher decides: "Based on this query and this user's history, the correct route is Route_Discovery."

Synthetic

  1. Similar except you ask an LLM to generate example search queries and corresponding actions.

“We then prompt a powerful LLM … using a carefully crafted prompt that includes task-specific instructions and manually curated few-shot examples.”

What do they mean by “task-specific instructions”?

  1. Intent Classification: Instructions to distinguish between "focused" intents (specific items) and "exploratory" intents (vibe, genre, or discovery)
  2. Hypothesis Generation: Guidance on how to handle ambiguity. If a query suggests "several plausible user intents," the instruction is to select a route that triggers parallel tool calls to cover those hypotheses (e.g., the Lady Gaga "latest albums" example).
    1. Prioritization Logic: Instructions on which route to prefer if multiple seem applicable (e.g., "Always prioritize the Personalized_Recs route if the user says 'for me'").
  3. Pre-Fusion Selection: Instructions to pick from predefined bundles of tools rather than creating a plan from scratch, ensuring the response is just a simple "Route ID".
    1. Route Descriptions: A comprehensive "menu" or registry where every Route ID is described by the specific intent it satisfies (e.g., "Use Route_Similarity if the user compares two artists").
    2. Sub-Agent Triggering: Specific criteria for when to route a query to a sub-agent like the AI DJ, particularly for "broad music searches”

“To encourage diverse outputs, we sample responses from the teacher model using a high temperature setting”

Why do they want to encourage diverse outputs?

You want to have different candidates to compare. The more candidates you have, the more likely that you can find at least one candidate that passes the criteria. Having higher variance in candidate quality is fine for now because rejection sampling will filter out the low-quality candidates.

How does Rejection sampling Fine-Tuning work? How does the reward model work?

RFT

  1. Diverse sampling from LLM with high temperature.
  2. Only candidates that meet the quality criteria of the reward model are given a "PASS" grade.
  3. The system takes the "winners" (the PASS candidates) and uses them as the "Gold Standard" training set. The smaller Student model is then fine-tuned to replicate the exact reasoning of those winning examples.

Reward Model

  1. The reward model is often another powerful LLM (or the Teacher itself) provided with a Rubric.
  2. The Rubric (Criteria): The model is given specific metrics to check, such as:
    1. Relevance: Did the routing decision actually match the user’s intent?
    2. Technical Validity: Is the generated JSON schema valid and can it be parsed by the downstream tool?
    3. Freshness/Diversity: If it was a discovery query, did the parameters include filters that would lead to diverse, new content rather than just the same popular hits?

What is the LLM router (prompting)? Why are they comparing it to LLM router (post-training)?

The LLM router (prompting) refers to the smaller, efficient "student" model used in a standard prompt-based setup without the specific post-training adaptation. In this mode, the model is simply given instructions and few-shot examples in its input context to perform the routing task, rather than being fine-tuned on synthetic data.

Spotify compares these two to demonstrate that fine-tuning (post-training) the smaller model on high-quality synthetic data is superior to just prompting it.

Why would there be an improvement in quality relative to the teacher model? Is it because of the RFT stage?

Yes. The teacher model isn't perfect. RFT uses an "LLM-as-a-judge" to discard the teacher's mistakes and hallucinations, ensuring the student is trained only on "PASS" (gold-standard) examples.

By sampling multiple responses at high temperature, the system can find an "optimal" routing decision that the teacher might not have picked on its first try. The student then learns this superior logic.

While the teacher is a broad generalist, post-training turns the student into a narrow "expert." This removes the inconsistency found in large-scale models and results in more predictable, high-quality routing for this specific task.

“To evaluate the effectiveness of the system, we adopt an LLM-as-a judge approach. This involves using a powerful language model with access to real-time world knowledge”

How does it have access to real-time world knowledge?

  1. Metadata Enrichment from internal data sources that are updated regularly.
  2. Tool Use: frontier models (like GPT-4o or Gemini 1.5 Pro) often have Search Tooling enabled.

If they’re using a powerful model for LLM-as-a judge, could it end up being really slow? How do you fix that?

Each individual call might be slow but there is a very large scope for parallelization here since it’s all offline.

Is there any reason why “Finding similar artists” would improve a lot more than “Broad podcast search”?

The “Finding similar artists” use-case is very incompatible with traditional search architectures (no good way to use lexical or semantic similarity for this) but it can easily be handled by agentic approaches. “Broad podcast search” is already reasonably well-suited for traditional search architectures.

Other blog posts I’ve read focused on metrics such as tool recall and tool precision. Is Spotify right not to focus on those here?

Because of the PFR architecture, the relevant technical metric is Route Classification Accuracy, not tool recall and tool precision.

Discussing intermediate metrics like Route Classification Accuracy would be helpful for outlining how they would go about debugging this pipeline but doesn’t matter much for the end-user. Music discovery is about "vibes." There is no legal requirement to hit a specific similarity API if the final recommendation is something the user loves. Spotify can afford to prioritize the experience over the execution path

“substantial improvements in search across multiple use-cases”

All of the improvements they give are for specific subsets; should we be worried about the multiple testing problem?

Depends on how many such subsets they evaluated; the subsets cited in the paper are pretty broad though so there can’t have been too many other subsets.

It isn’t clear from the blogpost but these subsets may also have been pre-defined, which would help mitigate concerns about p-hacking.

If I had to prove to a business leader that all of the work building this was worth it, how would I do that?

  1. For each of the use-cases, you’d have to cite what proportion of searches these represented, not just the improvement within the use-case.
  2. Any latency and cost-savings?
  3. Reduced technical debt?

How would you improve this system if you were building it today?

  1. Transition from RFT to GRPO: Instead of a static "PASS/FAIL" from a teacher model, use Group-Relative Policy Optimization (GRPO). This allows the model to learn from "relative" quality (e.g., Route A is better than Route B) and directly optimizes for user CTR (Click-Through Rate) and Long-term Retention rather than just a judge’s opinion.
  2. Multi-Task Alignment: In 2026, we use multi-task learning to align instantaneous intent (Did they find the song?) with long-term happiness (Did they keep listening for 30 minutes?). This prevents the router from picking a "clickbait" route that provides a fast but unsatisfying result.
  3. Instead of exact-text matching, implement Semantic Caching
  4. Speculative Routing (Cascades): deploy a Speculative Cascade. A tiny, 100M-parameter "Draft Router" makes a split-second guess. If its confidence is >95%, the tools fire immediately. If not, the request "cascades" to the larger Student model. This shaves an additional 100–150ms off the p95 latency.

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Learn more about RFT
  2. Learn more about GRPO
  3. Learn more about Semantic Caching
  4. Learn more about evaluation strategies and metrics
2025 Blogpost original

How Agentic Search Unlocks Legal Research Intelligence at Harvey

A simple RAG approach consisting of one-shot retrieval cannot adapt its strategy as it discovers information. It can't decide which knowledge sources are relevant; and it can't determine when it has gathered sufficient context. Naive RAG…

My notes

Summary

Context

Harvey Assistant facilitates legal research by answering user queries through reasoning on relevant data sources and providing citations.

Problem

A simple RAG approach consisting of one-shot retrieval cannot adapt its strategy as it discovers information. It can't decide which knowledge sources are relevant; and it can't determine when it has gathered sufficient context. Naive RAG works for simple queries but breaks down when research requires multiple rounds of information gathering, synthesis across sources, or reasoning about missing information”

Solution

Agentic Search and Reasoning

What I didn’t understand or am unsure about

Question

Answer

Let’s say you try doing the due diligence task they mentioned in the blogpost using a naive RAG approach. What specific issues would result?

  1. Legal jargon may not be well-suited to the retrieval systems most standard RAG systems use. For example, non-solicitation of customers is closely related to non-competes but this sort of similarity may not be reflected in the embeddings used for retrieval.
  2. Similarly, employment agreements (private contractual language) and public statutory/case law have different semantic spaces (i.e. the embeddings of their chunks probably reside in different clusters) and the user query is probably more similar to one vs the other depending on the language. Because of this, most of the retrieved chunks will be from whichever document type was closer in this space.
  3. It could be helpful to have additional searches based on the initial retrieved context. For example, the retrieved context references another section or document whose details are relevant for the final response.
  4. If the retrieval step fails to find the exact statute, the Large Language Model (LLM) may use its internal "training data" to fill the gap. This could lead to hallucinations.
  5. Naive RAG suffers from "Single Query" limitations - it tries to find one set of chunks that satisfies all three needs simultaneously. It cannot use multiple different queries to retrieve different kinds of relevant context.

How does agentic search enable “Real-time relevance” as opposed to a non-agentic approach? The knowledge base would be fixed regardless so are we assuming that the agentic approach adds web-browsing capabilities?

Although the knowledge base consisting of a vector database of embedded text is fixed, agents have access to tools like web-browsing/targeted scrapers and API calling (to pull third party data).

It also doesn’t need to wait until the next indexing cycle; it can just retrieve new documents directly from the database if needed.

What’s an example of how the agentic search could rewrite or draft a new query to make it more suitable for retrieval?

Original user query: California non-compete 2024 updates

Agent rewritten query: California SB 699 OR AB 1076 non-compete enforceability 2024

These are two non-compete bills in CA that an agent can associate with the user’s query because of its pretraining knowledge.

What does it mean to “interleave reasoning traces with tool calls”?

ReAct (Reasoning + Acting) paradigm. Interleave just means to do these steps one after the other repeatedly. Reasoning traces are provided as context when determining what tool to call.

“The agent analyzes what information is needed and develops a search strategy.”

What is a search strategy? Is it a list of queries it wants to search?

Not just queries but data sources, sequencing, contingency planning, effort calibration (stop after one query and keep iterating until completeness check is satisfied).

“For complex queries, it identifies which knowledge sources will be relevant.”

Complex queries might have multiple components, each of which require different queries to retrieve relevant data sources. How could their approach handle this?

Fragmenting the search into source-specific sub-queries and then unifying the results through an iterative reasoning loop

  1. Dependency Graph. This identifies which parts of the request can be done at the same time (parallel) and which parts must wait for information from a previous step (sequential).
  2. Parameterized Tool Routing: each data source can have a different retrieval strategy

“Reasoning and Synthesis: The agent reasons about how the retrieved information connects and applies legal standards to specific facts or provisions.”

In this situation, is the agent’s LLM storing all of the retrieved information in its current context window and generating a response based on that?

If it's a simple task: Yes, it will likely pull all relevant chunks into the window and synthesize them in one go.

If it's "Deep Research": If the agent retrieves 500 pages of case law, it won't just paste it all. It will perform incremental synthesis. It reads a batch, summarizes the "legal rule" or "relevant fact" into the reasoning trace, and then may discard the raw text of those documents to keep the context window "clean" and focused on the high-level analysis.

“The agent evaluates whether it has sufficient information to fully address the query. If gaps remain, it returns to Step 2 for additional retrieval rounds with refined queries.”

At this point, the agent has already generated a response. When returning to Step 2, does it look at the newly retrieved context alongside the existing response and generate a new response based on that?

Yes (although to be precise, the agent hasn’t generated a response at Step 3 but an internal monologue containing a reasoning trace and its synthesis).

What exactly is the difference between Step 3 and Step 5 besides citations?

The output of Step 3 is more of an internal monologue than a final deliverable. In Step 3, the synthesis is used to "reason about how information connects" to find what is missing. In Step 5, the synthesis is used to "fully address the query" in a comprehensive manner.

Since Step 3 is being used only by the agent whereas Step 5’s output is shown to users, the kind of language use will probably be different.

How are the citations generated?

It’s not clear from the blogpost but we can assume from industry best practices that:

  1. Observation IDs: In Step 2 (Retrieval), every chunk of text the agent pulls is assigned a unique identifier (a "Pointer"). When that chunk is "observed" by the AI, its ID stays attached to it.
  2. The Chain of Custody: In Step 3 (Reasoning), when the AI says, "This clause is unenforceable," its internal logic window (the Reasoning Trace) includes a note: [Source: Vault_Doc_A, Page 4].
  3. Verification Agents: As mentioned in Harvey's February 2026 technical updates, they utilize a specialized Citation Quality Agent. This is a secondary, background process that "fact-checks" the final response. It looks at the drafted sentence and the cited document side-by-side to ensure the document actually supports the legal claim being made.

What is the difference on the backend between Deep Research and regular queries? In other words, if a user selects Deep Research, what specific agent configurations are modified to reflect this selection?

  1. Agents are often configured with a number of iterations budget for retrievals; deep research probably has higher allowed iterations.
  2. The completeness check in Step 4 is probably more rigorous.
  3. More complicated (but slower) model

For the example they gave with a lawyer asking “What's our exposure on the termination provisions” without explicitly enumerating relevant sources, how can the agent actually determine what’s needed?

  1. Each tool the agent has access to has docstrings which can help connect the query to the agent’s capabilities.
  2. The system prompt might have instructions or examples. We could tell it to err on the side of over-retrieving.

“The agent lacked a calibrated sense of when "good enough" was sufficient and when a query demanded exhaustive coverage.”

How can you help the agent fix this gap?

  1. Complexity Classifiers (The "Router"): Before any research begins, a lightweight "Intent Model" classifies the query on a scale of 1 to 5.
  • Level 1 (Direct Retrieval): Minimal effort, low token budget.
  • Level 5 (Deep Research): Unlocks the full iterative loop and allows for parallel searching.
  1. Instructions or examples in the system prompt
  2. Stopping criteria/checklist
  3. DPO

Why is going from single-source to multi-source citations difficult?

  1. The LLM might lose track of the different data sources and mix them up. For example, if it stored a citation as section 5.1.1 or page 7 but forgot to include the reference to the specific source, it’s no longer clear what that citation is referring to.

“After receiving aggregated feedback and support tickets that describe issues …”

Is there any significant concern that this data is unreflective of actual user usage? Could a lack of positive examples be a problem?

Support tickets are inherently a collection of "Hard Negatives"—cases where the AI failed. If you only optimize based on these, you risk Overfitting to edge cases and degrade the AI’s performance on common, simple queries that were working well previously.

This can be mitigated in the Expert Query Generation by creating the set of evaluation queries based on the real usage patterns distribution.

“metrics like hallucination, tool recall, retrieval recall, formatting, and answer quality”

How are all of these metrics defined and measured?

  1. Hallucination: degree to which the agent’s claims are not supported by the retrieved context
    1. LLM as a judge is given the final answer and the retrieved chunks to determine if each sentence in the answer is "entailed" by the context.
  2. Tool Recall: this measures the agent’s ability to select the correct tool from its toolbox
    1. Recall = Correct tools called/total tools required for task
    2. Labeled data generated by reverse-engineering tool call scenarios. Use a tool for something and then create a query that can only be answered correctly using that tool.
  3. Retrieval Recall: measures if the tool actually found the right information inside the database
    1. Recall@K: % of all relevant documents that were present in the retrieved set
    2. Labeled data obtained similarly to tool recall. Reverse-engineering queries from docs.
  4. Formatting: measures the agent's adherence to structural and stylistic constraints
    1. Heuristic "validators."
  5. Answer Quality
    1. LLM as a judge?

“We iteratively refine tool design, system prompts, tool bundles and preambles, and tool docstrings.”

What do they mean by refining tool design? What are tool bundles?

  1. Refining tool design might mean making the code more AI-friendly, such as by reducing the number of possible parameters and changing the output format.
  2. Logical groupings of these tools; the system doesn't give the AI every tool for every query; it gives it a "bundle" relevant to the task.

“This approach solved our source selection challenge by giving the agent clearer signals about when to use each knowledge source — improving tool selection precision from near zero to 0.8-0.9.”

What does knowing how to use each knowledge source have to do with tool selection precision?

Also, how are they measuring tool selection precision? Where did the data for that come from?

The tools they’re referring to might be stuff like “search LexisNexis”; in most use-cases, tools means something much broader but I guess here they just mean Retrieval Tools, not Functional Tools.

Tool precision is probably measured as correct tools called/total tools called, with the labeled data coming in a similar way as the tool recall described above.

How would you improve this system if you were building it today?

It’s hard to know for sure because some of their descriptions lack specifics but ideas include:

  1. Using an adversarial Prosecutor Agent for the Completeness Check in Step 4 that acts as opposing counsel. This can help stress-test the model’s response and avoid confirmation bias.
  2. Incorporate other user’s conversations into the retrievable database; if another user at the same firm asked a similar question, the agent should be able to use that as a reference.
  3. Dynamic Decision Checkpoints: the user can provide input mid-response based on the reasoning trace so far to guide the agent in the right direction
2025 Paper original

Improving Pinterest Search Relevance Using Large Language Models

Improving relevance of search results

My notes

Summary

Context

Users can search for content on Pinterest; they type in a query and either press enter on their own query or select a recommendation. Then, a bunch of pins pop up.

Figure from the post

Problem

Improving relevance of search results

Solution

Fine-tune LLM on human labeled search relevance data; distill LLM into a smaller model for serving.

Some key insights/new things I learned

  1. Data Labeling:
    1. Humans: human-labeled data focused on US queries, rating queries on a scale from 1 to 5 for relevance
    2. Fine-tuned LLM: LLM fine-tuned using human labeled data, then used to label all languages and countries
  2. Features:
    1. Teacher: Pin titles and descriptions, Synthetic image captions, High-engagement query tokens, User-curated board titles, Link titles and descriptions
    2. Student: Query-level features (query interest features, shopping interest features, and SearchSAGE query embeddings), Pin-level features (PinSAGE embedding, visual embedding for the image, and SearchSAGE Pin embeddings) and Query-Pin interaction features (BM25 and text match scores for different text fields, historical engagement rates between the Pin and query, etc)
  3. Architecture:
    1. Teacher: Cross-encoder language model
    2. Student: MLP
    3. Distillation: Student trained on labels generated by teacher.
  4. Serving: {}

What I didn’t understand or am unsure about

Question

Answer

How is a relevance objective different from predicting engagement? Is it because the labels are human reviewer labels instead of CTR data?

Also, why do we want to prioritize relevance (defined in this way) over engagement?

The blog explicitly states that the relevance model is trained using human-annotated data based on a 5-level guideline (ranging from L1 to L5). This is contrasted with engagement data, which the blog refers to as "past user engagement" or "logged search engagement and impression dataset."

A user might click pins because they seem attractive even if it doesn’t actually answer their query (clickbait). Prioritizing engagement over relevance may increase the likelihood of the user clicking the pin but may detract from the long-term user experience.

On a more general note, a purely engagement-based objective could prevent new trends from surfacing to users.

What are “query interest” and “shopping interest” features?

This isn’t clear from the blogpost but Gemini suggested that

  • Query Interest Features: These likely represent the taxonomy or category mapping of a search term.
    • Example: If a user searches for "Mid-century modern side table," the "query interest" feature might map this to broader categories like "Home Decor" or "Furniture." This helps the model understand the high-level "bucket" the query belongs to, which helps it stay within the right neighborhood of content even if the specific words don't match perfectly.
  • Shopping Interest Features: These likely represent the commercial intent or product-specific metadata associated with a query.
    • Example: A search for "red velvet cake recipes" has high informational intent, while "red velvet curtains" has high shopping intent. These features likely signal to the model whether the user is looking for a product to buy or an idea to save. It might also involve mapping the query to a "Product Type" taxonomy used in Pinterest’s shopping catalog.

“owing to the seasonality in Pinterest Search. By using a multilingual LLM-based teacher model, we are able to successfully generalize from human-labeled data focused on US queries to unseen languages and countries.”

What seasonality are they talking about here and why is it relevant? How does using a multilingual LLM-based teacher model let them generalize?

Users may search for slightly different things depending on the time of year (for example, searching for Christmas related stuff in December). Depending on how the annotated data was collected, it may not contain a representative sample of user search queries.

To tackle seasonality, the LLM teacher model can leverage its preexisting knowledge of the world.

For new languages, LLM teacher model can be trained on this data but still be able to generalize to other languages through Cross-Lingual Transfer. Multilingual LLMs (like XLM-R) are trained so that words with the same meaning in different languages are mapped to similar mathematical "vectors".

They fine-tuned the teacher LLM on human labels from US queries and then extended it to other languages. They may show in the evaluation section that this worked well but before starting, could they have safely assumed that this approach would succeed?

In the AI research community, it is a well-documented property of multilingual models (like XLM-R) that they can perform "cross-lingual transfer."

There is a risk that this doesn’t apply the same way to Pins on Pinterest but small enough to justify trying it regardless.

How does a cross-encoder language model like the one described in the blogpost work? How is a cross-encoder different from a regular encoder?

Cross-encoders like the one in the blogpost work by

  • The query and the Pin text are concatenated together into a single sequence (e.g., [CLS] query [SEP] pin_text [SEP])
  • This combined text is fed into a transformer model
  • For Encoder Models (mBERT, mDeBERTa, XLM-RoBERTa)
    • To get the 1–5 score, Pinterest could take the vector representing the [CLS] token from the final layer and pass it through a Linear Layer (a simple neural network layer) followed by a Softmax function. This outputs the probability for each of the five classes (L1, L2, L3, L4, L5).
  • For Decoder Models (Llama-3-8B)
    • Instead of the first token ([CLS]), they use the final hidden state of the very last token in the sequence (often the [EOS] or End-of-Sentence token). In a decoder model, the last token is the only one that has "seen" all the previous tokens in the sequence. OR
    • They might designate a specific padding token as the classification head

This is different from a regular encoder because

  • A regular encoder encodes the Query and the Pin separately whereas a cross-encoder encodes the Query and the Pin together in a single pass. The cross-encoder approach allows for Full Self-Attention. Every word in the query can "attend" to every word in the Pin text.
  • Regular encoders are generally faster because you can pre-calculate the Pin vectors and store them in a database. However, it is less accurate because the model can't see how specific words in the query relate to specific words in the Pin.

Does the teacher model use ordinal classification for the relevance score prediction or just regular classification? Does it matter?

It doesn’t seem like it uses ordinal regression but I’ve rarely seen it in industry anyway so I don’t think it matters too much in terms of final model performance. Standard LLM fine-tuning libraries are natively built for cross-entropy so sticking with regular multiclass classification makes life easier.

Why doesn’t the teacher model use SearchSAGE or PinSAGE or any of the other features listed for the student model?

  • {I’ll need to double check this but }SearchSAGE and PinSAGE might be trained using engagement data whereas they want to focus on relevance for the teacher model.
  • The SearchSAGE and PinSAGE embeddings might also not be compatible with the cross-encoder architecture they describe for the teacher model.
  • The raw text is probably more informative than the embeddings anyway.

How much slower is the cross-encoder based teacher vs the MLP based student model?

Gemini’s best guess is

  • (Teacher) Cross-encoder: 12ms per query-Pin pair
  • (Student) MLP: 0.01ms to 0.1ms per query-Pin pair

What exactly is the architecture of the student model?

Multi-Layer Perceptron, probably with 3 to 5 layers with decreasing widths and a ReLU (Rectified Linear Unit) activation function to introduce non-linearity. The final layer would be a Softmax layer with 5 nodes.

The student model architecture seems relatively simple - how does it end up working so well?

  • It’s using embeddings that have already done a lot of the heavy lifting of understanding queries and pins
  • It’s trained on a ton of data
  • The hand-engineered features like BM25 are pretty powerful on their own too/

Does the student model use the human annotated data directly at all or is only trained using the teacher model’s labels?

The blogpost implies that it only uses the teacher-generated labels but industry best practices often mix both datasets.

Did they make any changes to their model distillation vs industry-standard approaches?

Standard model distillation teaches the student to mimic the probability distribution of the teacher, not just the labels (which it seems is what they’ve done).

How would you go about determining the appropriate latency-performance balance?

The 100ms Rule by Google: “user interface should respond to user input within 100 milliseconds to feel instantaneous”. Look at P99 latency, accounting for network time and database lookups.

Is there a good theoretical reason why the AUROCs for the 3+ threshold are higher than for 4+ or 5+ in this table?

Maybe less positive examples for higher thresholds?

“sequential addition of these text features”

Is there any reason they chose the sequence they did? Does it matter?

Seems like they followed an Internal to External and High Coverage to Lower Coverage order.

The order matters in terms of demonstrating individual feature importance but not for comparing the effect of the presence of the full set to the absence of the full set.

“Table 5 demonstrates the benefits of training on increasing amounts of augmented teacher-generated labels.”

What does the word augmented mean here?

I think they just mean the data they generated using the LLM which in some sense “augmented” the human data.

I guess they stopped at 30M distilled labels when evaluating based on “Table 5: Comparisons of production model performance when training on different amounts of labels”. If I'm a decision maker tasked with determining how many labels I should generate, how do I make that determination prior to starting the process? I'm assuming that once I've started, I can just look for performance plateaus (although is there a more systematic way to do that part too?).

You might be able to predict the "right" number of labels before spending your entire compute budget by using Scaling Laws.

You can also calculate the marginal Cost per 0.1% Gain and stop when it stops being worth it.

Is nDCG@K considered the best search metric? What are the tradeoffs of using it vs other metrics?

It is widely considered the industry standard for evaluating search relevance when you have graded labels. It recognizes that a "Perfect" (L5) match is better than a "Good" (L4) match. Binary metrics like Precision or Recall treat them as exactly the same. It understands that a relevant result at position #1 is much more valuable than the same result at position #10. It "discounts" the value of a result as it moves further down the list.

The main disadvantage is that it’s mathematically complex; harder to explain to non-technical stakeholders.

How do people generally pick the k in the nDCG@k? How would Pinterest have picked 20 here?

In search engineering, k is almost always chosen to match the "effective session depth"—how far down the page a typical user actually goes before they either find what they want or give up.

On Pinterest, my guess is that the typical user looks through a lot more pins than the number of search results they’d scroll through on Google.

“This metric is defined as the number of search sessions that result in a high-significance user action.”

What's a high-significance user action?

Maybe a save, Closeup or Long-dwell Closeup.

How would you improve this system if you were building it today?

  1. Native Multimodal Teachers - Vision-Language Model (VLM)
  2. Upgrade the Student: From MLP to Small Language Models (SLMs)
  3. "Soft-Target" Distillation (Dark Knowledge)
    1. Full Distribution Distillation; teacher provides the Logits
  4. Active Learning & "Uncertainty Sampling"
    1. Have the Student model "flag" the Pins it is most confused about. The expensive Teacher LLM only labels those "High-Uncertainty" cases.
    2. {ADD MORE DETAILS ABOUT HOW THIS WOULD WORK}

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Learn more about cross-encoders
  2. Learn more about Wide & Deep neural networks
2024 Blogpost original

Building Mature Content Detection for Mod Tools

Mature Content filters (MCFs) to automatically filter NSFW content (e.g. sexual and graphic images/videos)

My notes

Summary

Context

Mature Content filters (MCFs) to automatically filter NSFW content (e.g. sexual and graphic images/videos)

Solution

Machine Learning model

Some key insights/new things I learned

  1. Data Labeling: created detailed classification guidelines according to content policy, and had a production dataset labeled with the classification
  2. Features: “We found that a mix of visual signals, post-level and subreddit-level signals as inputs produced the best image and video classification models.”
  3. Architecture: simple v0 heuristic model, then v1 model, then more advanced deep learning models
  4. Serving: added it to Gazette, Reddit’s internal ML inference service. During user media upload, Reddit’s Media-service notifies Content Classification Service (CCS). CCS, a main backend service owned by Safety for content classification, collects different levels of signals of images/videos in real-time, and sends the assembled feature vector to our safety moderation models hosted by Gazette to conduct online inference.

What I didn’t understand or am unsure about

Question

Answer

What is the difference between their system and the Content Classification Service they mention?

CCS assembles the feature vectors (the inputs) and routes the data. Their system runs the newly developed models to provide classifications for the MCF.

How do you generally determine how much labeled data you need? In this case, how could they have done it?

Depends on what model architecture you want to use. For deep learning, more labeled data is typically needed. I found some information online suggesting the amount of labeled data should be 10-20x the number of parameters. Other places say 1000 labeled examples is a good place to start.

One other consideration is how prevalent the violative content is. Ideally, you’d want there to be enough data labeled as violative for the model to understand what factors make something violative. If the general distribution contains very little violative content, (like <0.1%), we might either need more data or some more targeted labeling (like looking at certain subreddits we think are more likely to contain violative content).

Once the initial labeled data is collected, you can plot the model's performance against the amount of training data and see if it’s plateauing. If it’s not plateauing, you should collect more data.

Also, look at subcategories to identify specific areas of underperformance and collect more labeled data within those subcategories. Likewise, focus on labeling data where the model was less sure (high entropy).

They mention “several iterations of annotation” so it looks like they’ve sorta done this.

“During user media upload, Reddit’s Media-service notifies … CCS {which} collects different levels of signals of images/videos in real-time…”

Users can upload videos of length up to 15 minutes to Reddit. How can they do real-time inference and still have good model performance?

  1. Use metadata where it’s informative enough to provide a confident classification
  2. Split the video into smaller "chunks" or "segments" that are processed by multiple workers simultaneously in the background
  3. Temporal sampling—extracting one frame every X seconds
  4. Caching (for exact replicas and clips/other small modifications)
    1. Segment Hashing

What’s our best guess for the model architectures of the:

  1. simple v0 heuristic model,
  2. v1 model
  3. more advanced deep learning models
  1. RegEx (Regular Expressions) for text, Subreddit Whitelists/Blacklists, and Perceptual Hashing (pHash/dHash) for media
  2. Gradient Boosted Decision Tree with pre-trained Convolutional Neural Network for image embedding, and post-level features (word counts, user account age) and subreddit-level features
  3. Vision Transformers (ViT) or CLIP-style (Contrastive Language-Image Pre-training)
    1. Video models could be VideoMAE or TimeSformer

“we needed to generate test images and videos that, while not inherently explicit or violent, would deliberately yield positive model classifications when processed by our system”

What does this mean?

To test the system, they wanted to send positive examples without risking employee wellness from looking at violative content and aligning with compliance guidelines.

The easy way to do this is to max out the metadata feature values for predictive features. You can also, given the model weights, reverse engineer random positive examples that don't look like anything to humans (Activation Maximization/Gradient Ascent).

“our SLA is to guarantee near real-time content detection”

What does real-time actually mean?

Consider the gap between user upload being completely processed and visible, and the first n views. This is measurable using historical data.

Then decide some cutoff, like “The classification must be completed within the time between a post being successfully uploaded and its first 10 organic views”.

Also measure the latency of all other parts of the process to determine how much leeway you have with the model.

“We've implemented multiple measures to ensure that our services can automatically and effectively scale during upstream incidents and periods of high volume.”

What would be some examples of such measures?

  1. Instead of trying to process every image the exact millisecond it is uploaded, Reddit pushes the task into a queue. During a "period of high volume," the queue acts as a buffer. If 10,000 images are uploaded in one second, they sit in the queue while the ML workers (Gazette) process them as fast as possible
  2. Separating image and video processing. If there is a spike in short-form image memes, they can spin up 100 extra image-processing instances without wasting expensive GPU resources meant for heavy video files.
  3. Caching "Features" to Skip Inference
  4. If the ML models are taking too long or the system is overwhelmed, the "Circuit Breaker" might temporarily switch to a "Safe Mode." For example, it might default to filtering anything from a high-risk subreddit into the modqueue without waiting for the ML result, ensuring that safety is maintained even if the tech is struggling.

“During production data sampling, having data annotated by our third-party annotation platform, automatically generating model metrics to gauge model performance over time”

What does this mean?

  1. Instead of just testing the model on old "historical" data, the system takes a small, random slice (e.g., 1%) of live images and videos currently being uploaded to Reddit today; send that sampled 1% of live data to a third-party platform.
  2. Model Drift.
    1. Data Drift: The type of content changes. For example, if people start using new slang or visual filters to hide "V" (Violent) content, the model’s performance will drop.
    2. Concept Drift: Reddit’s policies might change. If a new policy says "X" type of content is now allowed, the old model is suddenly "wrong" even though its code hasn't changed.
  3. Future" Automated Way: The system could do, if accuracy drops below a certain threshold (e.g., 90%), trigger the "automated model re-training pipelines" mentioned in the post.

What do they mean by “feedbacks of Mod ML models”?

MCF flagged content goes into community’s moderation queues where volunteer moderators review it; based on their feedback, the model can be automatically retrained.

What would the automated model re-training pipelines with active learning look like?

Uncertainty Sampling and Moderator Disagreement prioritized for retraining. Once the high-value "Hard Examples" are identified, they are sent to the third-party annotation platform mentioned in the blog. The pipeline could be triggered by the "automated model quality monitoring framework" when it detects performance drift or when enough new "hard labels" have been collected.

You can run the model in shadow mode and collect data on how it performs.

“In particular, we’re actively working with Reddit’s central ML org to support large model serving via GPU”

Does this mean the model they’re currently using was CPU-based? And what does that imply about the architecture (about which we don’t have too many other details)?

Might be two-stage: offline feature extraction with a heavy model and Online Shallow Classifier. Vision Transformers (ViT) or CLIP-style models would need GPUs.

They don’t mention the evaluation metrics they’re using but what should they be using?

Precision and Recall (heavy class imbalance so accuracy is not very useful); PR-AUC

Operational metrics: Latency, Throughput, Cost per Inference

What would be the best way to approach the overall problem outlined in the blogpost today? What improvements could they make?

  1. The cold start problem can be addressed by synthetic data generated by a diffusion model.
  2. Model Cascade: use a smaller model for near real-time classification and when the smaller model is unsure, send content to larger model (more parameters, reasoning. Can also use the pseudo labels for retraining. Experiment with different thresholds for uncertainty level to evaluate latency and cost vs quality.
  3. Agentic approach: context agent can fetch subreddit-specific context to guide decision making, adversarial audit agent can try to mitigate effects of adversarial noise added to content through relevant tool-use, orchestrator to manage these agents and the model cascade.

Level of understanding: 🤓

2024 Blogpost original

Innovating Email Protection: Writing Detection Rules with LLMs

My notes

Question

Answer

“Our rule engine is a DSL (domain specific language) that can express certain combinations of attributes that should be treated as malicious, such as a bad domain or malicious link”

How do contradictory rules get resolved?

[Post] Not addressed. The post shows one rule as a conjunction of attribute predicates and never discusses precedence or conflict.

[Inferred] The DSL structurally cannot produce logical contradictions. Every rule is a positive matcher — it fires (malicious) or it doesn't. No rule asserts "safe." So aggregation is a disjunction, and "contradiction" collapses into "one rule fired, others didn't," which is just a hit.

The real conflict is rule vs. model: a rule fires on a message the behavioral engine scores as clean.

[Search] Their published architecture answers that. Two documented combination patterns — "Ensemble of Models-as-Features: Each sub-model is a feature" and "Ensemble of Classifiers" — and they say they use "a combination of all the above approaches." So a rule hit is either a feature into the final classifier or an OR-ed override, depending on the path.

[Inferred] Actual operational conflicts get resolved by suppression precedence: safe-sender lists, per-tenant exclusions, rule kill-switches. Post-delivery architecture also permits retraction — disable a misfiring rule and restore what it quarantined.

Reframe worth having ready: in a positive-only DSL, conflict resolution is a precision-budgeting problem, not a logic problem. Every rule consumes part of a global false-positive budget. [Search] That budget is published: "The false-positive rate needs to be as low as one in a million." So the question becomes "how do you attribute FP budget across N rules and the model," which is what the backtesting infrastructure exists to answer.

“"Select a set of attacks that we want to detect. Study these attacks and craft a rule that matches them."”

What does this offer beyond what a model is capable of? Shouldn’t a model learn those rules automatically?

[Post] Three reasons given for keeping rules:

  • Interpretability — "deep neural networks are far more difficult to interpret than rules"
  • Speed on knowns — rules "quickly detect messages with known indicators of compromise"
  • Defense in depth — rules "augment a behavioral AI engine"

[Search] Two stronger reasons from elsewhere:

  • Cost. Shiebler: running LLMs over all mail is "Cost prohibitive." Scoring runs at "latencies of less than a second" over "terabytes of data a day." A DSL rule is an attribute lookup — effectively free.
  • Retraining lag — the real argument. Their own post: "adding new [features] involves retraining the model, which takes weeks." A rule ships in hours. That gap is why the layer exists.

On "shouldn't a model learn these automatically" — three reasons it doesn't, in principle:

  1. Sample efficiency. A new campaign has maybe dozens of examples. Gradient descent won't move a deep net on 50 examples out of billions without overfitting. A rule is a hard constraint writable from 5 examples — a hand-injected prior.
  2. Non-stationarity. The model is trained on the past; the attack is new. This is distribution shift, not a capacity problem. No amount of model quality fixes it.
  3. Controllability. You can't tell a neural net "stop flagging this vendor domain" and verify it happened. You can with a rule — auditable, revertible, unit-testable.

[Inferred] Honest counter-argument to have ready: everything a rule does could be done by a feature plus a fast-updating linear layer over frozen embeddings — same speed, same interpretability, better generalization. A rule is arguably that with a step function instead of a weight. The real reason to prefer rules is human auditability, not modeling power.

“Here is the set of malicious email messages that your rule should flag”

How is this set chosen? Manually or automatically? What level of granularity?

[Post] Not specified. It says sourcing attacks "can be fully automated," and lists "Selection of attacks to target" as decision point #1 — currently handled by vanilla software, with LLM involvement flagged as "an area of exploration in the future."

[Search] Their 2026 successor post (How Abnormal Taught AI Agents To Write Detectors) names the real trigger: Detection 360 customer submissions of missed attacks, going "from a customer-reported miss to a deployed detector in hours." Other plausible sources documented elsewhere: end-user reports, threat intelligence, and messages the ensemble scored just under threshold.

On granularity — [Inferred], but tightly constrained by the post's own framing. It calls the attack set "training data" and the rule the "trained model." That forces campaign-level granularity:

  • Single message → the LLM writes what amounts to a hash of that message. Zero generalization.
  • Attack type ("all BEC") → too heterogeneous. The only rule matching all of them also matches safe mail.
  • Campaign (a cluster sharing sender infrastructure, template, lure structure) → a genuine shared signature exists to extract. This is the only level where a one-shot translation prompt has a well-posed target.

[Inferred] Likely pipeline: cluster recent confirmed attacks (embedding similarity plus shared indicators), treat each cluster as one training set, generate one rule per cluster.

What features of the email message are passed into the LLM prompt?

[Post] Not stated — the prompt sketch shows only "Here is the set of malicious email messages" followed by an ellipsis.

[Post] Two constraints bound the answer:

  • "we extract thousands of rich signals describing everything about the message—content, sender, recipient, links, attachments, contextual information, and more"
  • The example rule mixes four kinds: a behavioral aggregate (never_seen_sender), registration metadata (from_fqdn_age_in_days), structure (attachment_extensions), and content (body_text_contents)

[Inferred] You cannot pass thousands of signals × N messages — that is a context and cost problem, and a signal-to-noise problem. So the prompt almost certainly holds:

  • The DSL schema — attribute names, types, operators ([Post]: "Here is the syntax of your rule engine")
  • A projection of each message onto that schema — attribute values, not raw email
  • Probably raw subject/body text too, since body_text_contents contains "urgent_language" implies content reasoning

[Inferred] The subtlety worth raising: the LLM can only write rules over attributes it is shown. Attribute selection is itself a retrieval problem — which of thousands of signals to surface for this cluster. Selecting by discriminativeness (where the cluster's distribution diverges most from a background sample) turns prompt construction into feature selection, and that is probably where most of the quality lives.

[Search] Their 2026 agent post confirms the direction — the agent now also receives "a random sample of normal traffic for the recipient." Background context was added later, which is exactly the gap the 2024 post admits to.

No calibration?

[Post] Correct — none, and the post is explicit about why: "the training data does not contain any safe messages. The algorithm relies on the LLM's prior knowledge of what safe business communication looks like." That is a stated substitution of an LLM prior for a measured negative distribution.

[Post] There is validation, just not calibration: step 4 is "Validate that the rule does not flag any safe messages," done by deterministic software after generation. FP control is a post-hoc filter, not part of fitting.

[Inferred] Why calibration is arguably the wrong frame: a DSL rule outputs a Boolean, not a probability, and calibration is a property of probabilistic scores. What you actually want per rule is:

  • a measured precision estimate on held-out traffic, and
  • a firing-rate estimate, since volume × (1 − precision) is the FP cost

[Inferred] The reframe: you don't calibrate the rule, you calibrate the decision to ship the rule. Backtest over a historical window, measure hits on labeled-safe traffic, set a launch threshold. If a rule should contribute to a score rather than force a verdict, learn its weight via the ensemble (Models-as-Features) — then calibration happens once at the ensemble output (Platt scaling, isotonic regression) instead of per rule.

[Search] They have the measurement half already: backtesting re-scores historical mail with point-in-time-correct features — "messages should be re-hydrated exactly as they would have been during online scoring."

Which model is used for this?

[Post] GPT-4 is named only illustratively: "large language models like GPT4 are excellent at this kind of translation." That is an example, not a production statement. The post also flags the alternative: "using a large language model that has been fine-tuned for security applications."

[Search] They have published that they do exactly that — "Fine-tuned 7B open-source LLMs (e.g., Mistral, LLAMA2)," trained with LoRA on "internally labeled email datasets," claiming 7B fine-tunes "can perform at the level of GPT-4" while being "significantly more cost-effective to run." Shiebler confirms: "We've utilized Falcon and LLaMA and fine-tuned those to a few different tasks."

[Search] Two caveats on attribution:

  • Those fine-tunes are described for detection/classification generally, not specifically for rule generation.
  • Rule generation is offline — one call per campaign, not per message — so the cost pressure that drives small fine-tunes doesn't apply. A frontier model is affordable at that call volume.

Any latency constraints?

[Post] Never raised. But there are two distinct budgets and they are very different.

Rule generation — offline. [Inferred] One LLM call per attack cluster, human-free but not user-facing. Budget is minutes. Not latency-constrained; throughput- and cost-constrained (calls/day × tokens).

Rule evaluation — online and tight. [Search] "All this must be done at latencies of less than a second," over "terabytes of data a day."

[Search] Important nuance: post-delivery means that sub-second budget is not a block-or-deliver deadline. It is the budget to remediate before a user clicks — messages "briefly exist in the mailstore before remediation—typically measured in seconds."

[Inferred] This is the architectural point of the whole post. A DSL rule is a conjunction of already-computed attribute lookups, so evaluating thousands of rules per message costs microseconds. That is why the rule layer can be arbitrarily wide, and why "LLM writes the rule offline, cheap deterministic engine evaluates it online" is the trick: you pay LLM cost once per campaign, not once per message.

“rule evaluation infrastructure”

How are rules evaluated?

[Post] Named, never described — "invested heavily in rule evaluation infrastructure." Only the two criteria are given: matches the selected attacks, does not flag safe messages.

[Search] Their Re-Scoring series describes the infrastructure directly:

  • Purpose: "we measure what percentage of historical attacks we would catch today with the current detection system"
  • Scope: replays "an entire stack that includes multiple models, decision logic, feature extraction code, and datasets" — not just a model artifact
  • The hard part, point-in-time correctness: "messages should be re-hydrated exactly as they would have been during online scoring." Aggregates must use the historical window: "look up the count starting from 1 week ago and going 30 days back"
  • Execution: Spark over messages and labels copied from production into their data lake

[Search] The 2026 agent post adds a two-stage gate: "a fast local evaluation" against the attack plus a random sample of normal traffic, then "a much broader evaluation across a larger slice of the customer's real traffic." False-positive triage is automated — "we LLM-label each one to determine what it actually is."

[Inferred] Why point-in-time correctness matters specifically here: the example rule uses never_seen_sender and from_fqdn_age_in_days. Both are functions of when you ask. Evaluate that rule today against a six-month-old message and never_seen_sender is now false, and the domain is 180 days older. Without time travel on the feature store, backtesting systematically understates recall and overstates precision.

[Inferred] What step 4 quietly requires: "does not flag any safe messages" needs a large negative corpus. At a 1-in-a-million FP target, you need on the order of tens of millions of safe messages to distinguish "zero hits" from "fires once a day in production."

“For example, we could imagine building a system in which one LLM generates attacks and another writes rules to stop them.”

If we wanted to do this, how to make the attack generation LLM good?

[Inferred] The core problem: "good" here is not "realistic." A generator optimized for realism reproduces attacks you already catch. You want attacks that are (a) plausible enough to work on a human and (b) currently evasive. Only the second is useful, and they are different objectives.

Design, MECE by what each component supplies:

  1. Objective — evade the current detector, subject to a validity constraint. Score generated attacks with the production ensemble; reward low scores. Separately constrain validity: the message must still accomplish an attacker goal (credential harvest, wire transfer, callback). Without that constraint you get gibberish that scores low because it is out-of-distribution, not because it is evasive. Use a separate "does this still work as an attack?" judge as a hard filter.
  2. Grounding — seed from real attacks, mutate along attacker-controlled axes. The NLP post enumerates them for callback phishing: brands impersonated, receipt structure, phone-number presentation, placement of malicious text. Parameterize generation over exactly those degrees of freedom. Free-form generation drifts off-manifold.
  3. Diversity — explicit novelty pressure. Deduplicate by embedding distance from both existing attacks and each other, or the generator collapses onto one evasive template. Quality-diversity search (MAP-Elites over the attack axes) is the clean formulation.
  4. Curriculum — co-evolve, freezing one side at a time. Freeze the detector, let the generator find evasions; then freeze that attack set and update the detector. Simultaneous updates give standard GAN instability. Keep a replay buffer of old attacks in the eval set to catch regression.
  5. Reality anchor. Periodically check whether generated attacks resemble what actually arrives. If the generator's distribution diverges from real incoming attacks, you are hardening against an imaginary adversary. This is the failure mode that kills these projects.

[Inferred] Biggest risk to name explicitly: you are training your detector on synthetic negatives-of-itself — a closed loop. The generator inherits its own priors about what an email looks like, so you overfit to its blind spots and can lose real-world recall. Mitigation: synthetic attacks go in the training set only, never the evaluation set that gates launches. Keep the launch gate on real traffic.

[Search] Named resources to go deeper: Perez et al., Red Teaming Language Models with Language Models (2022) is the canonical automated red-teaming paper; the quality-diversity follow-up is Samvelyan et al., Rainbow Teaming (2024).

How would you improve this system if you were building it today?

MECE by pipeline stage.

1. Attack selection (step 1) — the least developed piece, by the post's own admission. Cluster confirmed attacks before generating: one rule per campaign, using embeddings plus shared infrastructure signals. Then rank clusters by expected value — volume × current miss rate × recipient risk. Today it is presumably FIFO off the Detection 360 queue.

2. Prompt construction — two changes.

  • Add negative examples. [Post] concedes the gap: "the training data does not contain any safe messages." Include safe mail from the same recipients and tenants. [Search] Their 2026 post did exactly this, which validates the critique.
  • Feature-select before prompting. Surface the attributes where the cluster diverges most from background, rather than dumping the schema. Makes prompt construction a statistical step, not a formatting step.

3. Generation — sample K rules, not one. A greedy decode gives no way to trade precision against recall. Generate a population, evaluate all, keep the Pareto front. LLM inference is cheap relative to the value of a good rule at this call volume.

4. Evaluation and gating.

  • Launch on precision with a confidence interval, not a point estimate — at a 1-in-a-million target, "zero FPs on 10,000 safe messages" tells you almost nothing
  • Require point-in-time-correct backtesting (the never_seen_sender problem)
  • Auto-canary: launch in score-contributing mode, measure live for a fixed window, then promote to verdict-forcing

5. Lifecycle — the biggest gap. [Post] describes launch and nothing after. Rules decay: the campaign ends, infrastructure rotates, and the rule becomes pure FP risk. Add per-rule monitoring (fire rate, precision, unique catches), automatic demotion when a rule stops catching, and expiry by default — a rule with no true positives in 30 days should have to re-justify itself.

6. Redundancy control. Automated generation grows rule count without bound, and rules overlap. Measure marginal catch — messages this rule catches that nothing else does — and prune zero-contribution rules. Otherwise you accumulate FP risk with no recall gain.

7. Structural change worth arguing for: emit rules as features, not verdicts. Let the ensemble learn their weights. [Search] They already have the pattern — "Ensemble of Models-as-Features: Each sub-model is a feature." This gets calibration for free (once, at the ensemble output), turns conflicts into a weighting problem instead of an override problem, and makes low-precision-but-informative rules usable instead of unshippable.

2025 Blogpost original

From RAG to Richness: How Ramp Revamped Industry Classification

Issues with the old system include: Categories were occasionally obviously incorrect Categories could be so generic it was unhelpful Categories for similar businesses could be different The system was not auditable or interpretable Also…

My notes

Summary

Context

Industry classification is important since Ramp has various services that are tailored to each customer based on their industry.

NAICS and SIC are the two standard taxonomies for industry classification but Ramp mainly used a third, homegrown, non-standard industry classification system.

Problem

Issues with the old system include:

  • Categories were occasionally obviously incorrect
  • Categories could be so generic it was unhelpful
  • Categories for similar businesses could be different
  • The system was not auditable or interpretable

Also translating from the homegrown Ramp system (as some teams needed to do) required complicated mappings.

Solution

They migrated all Ramp industry classification to NAICS codes. To enable this migration, they needed a classification model that could predict NAICS codes based on the information they had about the company. They developed a RAG pipeline to do this.

Some key insights/new things I learned

  1. RAG requires a query and knowledge base; the queries here are the businesses' information and the knowledge base is the NAICS codes.
  2. They evaluated the 2 stages (retrieval and final classification) separately
    1. acc@k for retrieval: how often is the correct NAICS code in the top k recommendations
    2. Fuzzy-accuracy for NAICS code selected by LLM
  3. For the final classification, they use a 2-prompt system
    1. 1st prompt: large number of recommendations generated in the retrieval stage fed to LLM but without specific descriptions. The LLM is tasked with returning a "small list of the most relevant codes."
    2. 2nd prompt: small list fed to LLM alongside context/descriptions and ask it to pick the best one

What I didn’t understand or am unsure about

Question

Answer

“Third-party solutions are a quick way to get good, general performance; however, Ramp has unique needs with complex data, and we decided to build an in-house model.”

It’s not discussed in the blogpost but what would be a good way to determine if exploring an in-house model is worthwhile?

It’s hard to make a fully objective decision here but some useful things to consider:

  1. Benchmark the best available third-party API against your own data; cost, latency, classification performance
  2. How much you care about having the flexibility to modify the model (not data-driven, judgement call)
  3. Are third-party solutions black-box in ways that matter to you?
  4. Opportunity cost; what are the best alternative uses for the resources (human and capital) that would be directed towards the in-house solution?
  5. Data privacy

What exactly do the queries look like here?

It’s supposed to consist of business attributes but it’s not clear what exactly the formatting is. One possible approach could be:

"Company: [Name] | Description: [Scraped Text] | Legacy Category: [Old Tag]"

What exactly does the knowledge base look like here?

Similar to the query, it’s not clear what exactly the formatting is. One possible example could be:

"NAICS Code: 561311 | Title: Employment Placement Agencies | Sector: Administrative and Support Services | Description: This industry comprises establishments primarily engaged in listing employment vacancies and in referring or placing applicants for employment."

This seems more like a search system with an LLM reranker than standard RAG. In particular, the generation step here is just asking the LLM to pick between the surfaced options. Am I missing something?

Also, couldn’t they have used a BERT-style classifier (DeBERTa-v3, ModernBERT) that would be cheaper/faster and potentially have better accuracy too?

Most RAG use-cases have less constrained generation than this blogpost but it’s still technically RAG.

They also “ask the LLM for justifications to clarify the reasoning behind each prediction” so there’s a little more to it.

Using encoder-only approaches for classification might require retraining the entire model if NAICS codes change {need to think more about this}. I think it might still have been worth trying though.

Would the best fuzzy-accuracy metric just be what % of NAICS code digits were correctly identified (Prefix-based) or could any kind of non-linearity/weighting be useful?

Can’t be sure without knowing more about how NAICS codes work but it might be helpful to give more weight to the left-most digits.

“If we naively optimize for acc@k we would end up with a system that just recommends the whole knowledge base (guaranteed that the correct label is present if we recommend all possible labels).”

  1. Since acc@k is for the top k positions, is this actually true?
  2. Could they have used mean reciprocal rank instead of acc@k to mitigate this concern as well?
  1. Not true for fixed k.
  2. They could’ve but since we still have an additional reranking stage, the MRR isn’t super interpretable.
  3. One interesting metric they could have used is precision@k which would penalize noise.

“We also identified economical embedding models”

What’s the original embedding model they are using (if specified) and what are more economical alternatives?

They don’t specify the original model but common choices are OpenAI text-embedding-3-large, Cohere embed-english-v3.0 and Google text-embedding-004.

Lighter weight versions include BGE-Small, GTE-Small and OpenAI text-embedding-3-small.

What do the different lines in the “Frequency Label in Top k recommendations” graph represent?

Different hyperparameter configurations for the retrieval stage.

The full version of this entails finding the Pareto Frontier and then satisficing (judgement call).

“We … selected those with the least resource requirements and best data coverage”

  1. What are the resources here?
  2. What’s meant by data coverage?
  1. Latency, $ cost, context size, memory footprint
  2. Some fields might be present for some businesses but not others. They prefer using those fields that are more frequently available.

“including more recommendations in the prompt gives the LLM a better chance at finding the correct code”

Is this necessarily true? Doesn’t the subsequent line “can lead to degraded performance if the LLM is unable to focus on the most relevant recommendations” contradict that?

P(Success) = P(Included) x P(Selected|Included)

It’s possible that P(Selected|Included) decreases faster than P(Included).

In the “Selecting Predictions” section, why is “Structured output class and fields” listed as a hyperparameter? What are the different choices here?

You can use Pydantic Models, the OpenAI API or other options. These structured outputs can constrain the output to not just be of a certain format, but also just to be selected from a specific list.

For the output, you can have it include a Confidence Score, Top-K Candidates, and other ancillary fields.

Is it actually sensible to evaluate the 2 stages separately if our final objective is only to predict the NAICS code correctly? Even if yes, should the retrieval objective function be given the same weight?

It’s useful for debugging when you’re trying to figure out if performance is held back by retrieval limitations or reranking limitations.

It’s not entirely clear how the final model selection was done though but my guess is that after filtering based on the acc@k from the first stage, they picked whatever had the best fuzzy accuracy.

“Likewise, longer or more descriptive information can help the LLM better understand a business or a NAICS code, but will also greatly increase the context size.”

They mention increased context size being a concern. Given the > 1 million size context windows that are common these days, to what extent is this still a fair concern?

Yes, price, latency, accuracy all could suffer from inflated context windows.

“we've also found cases where the LLM predicts the correct code despite it not being present in the recommendations”

How is this possible? Shouldn’t they constrain the output possibilities using structured outputs?

They probably didn’t implement structured outputs with list-constrained generation and in some cases, the LLM’s initial training was enough to figure it out.

What would be the best way to approach the overall problem they're trying to solve today?

  1. Assuming they used Bi-Enconders to generate the embeddings, late interaction models like ColBERT could be an upgrade.
  2. Agentic Routing instead of 2 prompts
  3. Small Language Models for the first prompt in the 2 prompt system
  4. Knowledge Graph
    1. If a certain code is retrieved in the retrieval stage, maybe it’s worthwhile to also retrieve other adjacent codes (similar to candidate expansion from “Building Airbnb categories”
  5. Dynamic Context Pruning
    1. Between the Reranker and the Final LLM Prompt

Level of understanding: 🤓

To-Do’s/Follow-ups

  1. Learn more about ColBERT/late interaction models.
2025 Blogpost original

Classifying Chaos: How Material Automates User-Reported Phishing at Scale

Security teams should prioritize the riskiest emails but user reports contain a mix of harmless, annoying and actually dangerous.

My notes

Summary

Context

At companies using Material, users can report suspicious emails to the security team.

Problem

Security teams should prioritize the riskiest emails but user reports contain a mix of harmless, annoying and actually dangerous.

Solution

ML-based User Report Auto Classification into safe, spam, or malicious categories.

Some key insights/new things I learned

  1. Some of the categories of features they used to predict include sender <> organization relationship, email content data, email metadata
  2. In addition to the classification, they provide explanations as well
  3. They have a manual classification mode as well that provides recommendations that the security team can then choose to accept or reject

What I didn’t understand or am unsure about

Question

Answer

How was the training and evaluation data for the models obtained? Is there any risk that it’s unrepresentative of the data they’d see in production?

This isn’t outlined in the blogpost.

The most likely source would be historical data from their own platform—specifically, email reports that were previously resolved by their customers' security teams.

If that is indeed the source of the data, the data distribution in training should reflect that seen in production, As time goes on, techniques may evolve and additional data may be needed to adjust for that.

They mention “analyzing thousands of signals” - how would they come up with thousands of signals? It must be some sort of automated process but what specifically?

They might mean embeddings (which have a high number of dimensions) as opposed to standard features.

Time-windowing for different features can also result in many more features.

If they have thousands of distinct features, should they just train a model on the full set of features or try to narrow it down? Would cost/latency be a concern if using the full set of features?

On a related note, is manual feature selection necessary if you’re using L1 regularization or an inherently regularized model like XGBoost or Random Forest?

They mention using “efficient architectures optimized for CPU inference”, suggesting they’ve reduced the complexity somehow.

Cost/latency could be a concern. One way this could happen is if we have a feature which doesn’t really contribute to the prediction but at inference time, requires making a slow API call or database query. Manual pruning using L1 regularization or tree-based importance could help prevent this unnecessary latency.

Dependng on how many observations we have, models may be inclined to find spurious patterns that do not generalize.

Different companies may have different definitions of safe, spam and malicious. Does Material address that in any way? If not, how could they have addressed it (if it’s even worth addressing)?

This should be accounted for to some extent through the sender <> organization relationship features.

On the broader question of different risk tolerances between companies, this is probably better handled at the workflow layer than the model layer. Companies could set different thresholds.

“we've also developed specialized embedding models trained specifically on email data”

How are these models trained? Are they fine-tuned with MLM (Masked Language Modeling) objectives or with classification objectives (requiring labeled data)?

Unclear from the blogpost but a reasonable approach could be to first fine-tune on email data with MLM (Domain Adaptation), then Task-Specific Fine-Tuning with Contrastive Learning for classification.

“we use an architecture that leverages subword information and hierarchical structure in text, allowing us to deploy these models effectively at scale”

What does this mean?

I think the subword information part is just standard subword tokenization (helpful for reducing vocabulary size).

The hierarchical part might be related to Hierarchical Attention Networks. You construct subword embeddings, word embeddings from the subword embeddings, sentence embeddings from the word embeddings, and then document embeddings from the sentence embeddings. Then feed that into a classification head.

How are the explanations for the classifications generated? LLMs looking at feature importance?

More likely Attention-Based Attribution with a Natural Language Template for the feature explanations.

“Classification consistency is maintained across related messages to prevent conflicting remediation actions - even when automated classifications are overridden by your team”

What does this mean?

If you make a decision about an email, that decision should affect the classification for other similar emails

“Case relationships are intelligently managed to ensure that different classifications for similar messages don't result in unintended consequences”

What does intelligent management look like in this context?

If one person marks an email as safe and another marks it as malicious, there should be a way to aggregate this information and make a joint decision.

What would the best approach to solve the overall problem the blogpost has outlined today?

  1. Maybe an agentic approach with a Multi-Agent Orchestrator.
    1. The Identity Agent: Checks the Relationship Graph (not just counts) to see if this "CEO" has ever messaged this "Intern" at 2 AM before.
    2. The Forensic Agent: Executes the URL in a headless "browser-agent" to see if it renders a polymorphic phishing page (detecting "Runtime Assembly" attacks).
    3. The Policy Agent: Consults the organization's specific security manual via RAG to see if this "Password Reset" follows internal protocol.
  2. Global Signals" to "User-Centric RAG
  3. Agentic Chain-of-Thought (CoT) logs for Explainability

Level of understanding: 🙂

2025 Blogpost original

Using AI to help SNAP recipients diagnose and restore lost benefits

The best way to resolve these interruptions isn’t always clear to users.

My notes

Summary

Context

SNAP benefits occasionally get interrupted for various reasons.

Problem

The best way to resolve these interruptions isn’t always clear to users.

Solution

  1. Built a webform that users can go through to figure out why their benefits were interrupted and how they can fix it.
  2. AI chatbot to help solve the same problem

Some key insights/new things I learned

  1. Decagon can be used to connect LLMs to internal knowledge bases and tooling to execute actions (agentic).

What I didn’t understand or am unsure about

Question

Answer

What is the difference between the two metrics mentioned in “Faster benefits restoration”?

The second one “Higher rates of restored benefits in the same month” calculates the rate of users prevented from churning and having to reapply. The first one “Days to next deposit” measures reductions in the length of benefits interruption (regardless of whether it was due to having to reapply or another kind of interruption).

Does the Decagon system have guardrails in place to prevent users from using the LLM for situations unrelated to the designated use-case?

Yes; see this blogpost.

What improvements could be made to the system described in the blogpost, given the state of LLMs and agents today?

  1. Image and voice support
  2. For any missing forms/documentation, have an agent actually fill out drafts

Level of understanding: 🤓

2025 Blogpost original

Using AI to help SNAP recipients make sense of “notices"

SNAP notices can be confusing to recipients for a variety of reasons.

My notes

Summary

Context

“Notices” are the formal communications that regulations require SNAP agencies to use when doing certain things that affect a person’s SNAP case.

Problem

SNAP notices can be confusing to recipients for a variety of reasons.

Solution

Use LLMs to help users determine (1) how important a notice is (2) what action is needed from the recipient

Some key insights/new things I learned

  1. AI seems like a promising approach to confusing SNAP notices because language friction is the main source of confusion and LLMs are pretty good at navigating that.
  2. One approach they might consider in the future to avoid hallucinations is chaining by having a separate model call evaluate the response to ensure all provided information is found on the original notice in some form.

What I didn’t understand or am unsure about

Question

Answer

What would be a good way to evaluate if the LLM is working well for this task?

  1. Build a ground truth reference set
    1. Gather a diverse set of notices annotated by Subject Matter Experts (SMEs) with The "True" Importance and the "True" Required Action
    2. Calculate a Risk-Weighted Confusion Matrix for the Importance and LLM-as-a-Judge for the Required Action (Faithfulness and Completeness)
  2. OCR Robustness Stress Test
    1. Add increasing levels of noise to images and see the breaking point of the LLM (does it start hallucinating at some point)
    2. You could also convert this to a score that can be used to trigger requests for photo retakes if the photo isn’t clear enough
  3. A/B test users on their understanding of the notices with and without the tool

Could it be helpful to use one-shot/few-shot prompting where we include examples of the kinds of outputs we want the tool to produce?

Probably, since the LLMs innate sense of what is important and unimportant may differ from the desired categorizations.

Also, a more advanced approach could be to use Dynamic Few-Shot where the examples added to the prompt are similar in some way to the user’s current notice.

What advantage does this system have over a user going to an LLM like ChatGPT directly and asking it about the notice?

  1. You don't know what you don't know. If a user thinks they understood the notice but they misinterpreted it, they’re not going to go to an LLM to ask for help. Propel’s system can be a useful default filter.
  2. Users may not know how best to prompt the LLM to extract the information that would be most useful.
  3. Security risks
  4. The best versions of LLMs are on the paid plans - the free tiers have much more limited file upload and language sophistication.

For the Q&A tool, why didn’t they implement it as a full chatbot where users can ask follow up questions if needed? Are there any risks/major downsides to this approach?

The output is more deterministic and there’s no risk that the model starts discussing areas outside of its area of expertise. You can also control the tone of the responses.

For a single response, you can implement Structured Outputs.

Is the notice analyzed in isolation or does it also have access to user characteristics that may affect the answer?

Seems to be analyzed in isolation - limits hallucinations and avoids introducing noise into the context window.

Level of understanding: 🤓

To-Do’s/Follow-ups

  1. Learn more about Dynamic Few-Shot
2024 Blogpost original

How Abnormal Security Leverages NLP to Thwart Cyberattacks

My notes

What I didn’t understand or am unsure about

Question

Answer

“we maintain models that use each character-level, subword-level, and phrase-level tokenization”

What does it mean for a system to use word-level or character-level tokenization? How does that translate to actual features?

Definition first: tokenization splits text into units and maps each unit to an integer ID. That ID indexes a lookup table of learned vectors (an embedding matrix). The tokenizer determines the vocabulary — the set of units the model can represent at all.

Concretely, "features" = a sequence of integer IDs, then a sequence of embedding vectors, i.e. a (length × embedding-dimension) matrix fed to the CNN. The model is not fed words; it is fed a matrix.

“Abnormal also utilizes neural networks to analyze all text fields in email messages, including the body, header, attachments, and links”

Separately or together?

[Inferred] Almost certainly both, at different layers — because the fields are distributionally different. A subject line, an SMTP header, a URL string, and body prose differ in length distribution, token distribution, and meaning. One shared encoder over a concatenation would have to learn all of them, and a long body would dilute a short decisive header signal. So: per-field encoders (sharing weights only where distributions are similar — definitely not URLs vs. prose), producing per-field embeddings or scores, fused downstream.

[Search] Their published architecture supports the fuse-downstream reading strongly:

  • "The core engine of detection is a multi-modal ML model" and "our ensemble of detection models use these data sources in various ways, and we have redundant detectors"
  • "Ensemble of Models-as-Features: Each sub-model is a feature" — per-field text model outputs become features to a final classifier
  • Features are "organized as a DAG of dependencies where some features rely on other features or models" — exactly what per-field-model-then-fuse looks like in a feature store

[Inferred] Best short answer: separately encoded, jointly decided. Text models run per field; their outputs join thousands of non-text signals in a feature DAG; a final ensemble produces the verdict. That is also how text evidence meets behavioral evidence — "sender never seen," "domain is 4 days old" — they land in the same feature space.

“Unlike the transformer, the CNN forces words that are close together in the text, such as words in the same or adjacent sentences, to be relevant to each other.”

What does this mean?

This describes a difference in inductive bias — the assumption baked into the architecture before any training happens.

Transformer / self-attention: every token can attend directly to every other token. The connection pattern is learned and all-to-all. [Post]: "Transformers connect each word in the text to every other word." Cost: attention is quadratic in sequence length — double the text, quadruple the work. [Post]: "this also makes them slow, due to the numerous connections they need to maintain."

CNN: a convolution slides a fixed-width filter — say 5 tokens — across the sequence. A token interacts only with tokens inside that window. Stacking layers grows the receptive field linearly with depth, so distant tokens interact only indirectly, and only if the network is deep enough. Cost: linear in sequence length.

So "forces words close together to be relevant to each other" means locality is hard-coded rather than learned. The model assumes the useful signal lives in local, n-gram-like patterns.

[Post] The payoff: "This simple assumption makes the model much faster to run. At Abnormal, we run CNNs on every message we process."

[Inferred] Two clarifications, because the sentence is slightly loose:

  1. "Forces… to be relevant" overstates it. The CNN does not force relevance; it restricts which interactions are directly computable. A filter can learn to ignore its neighbors entirely. The constraint is on connectivity, not importance.
  2. The trade has a real security cost. A CNN cannot easily represent a dependency between the first and last sentence of a long email — a benign opening and a payment-instruction close. [Post] mitigates by scope, not architecture: run text models on every field so no single long-range dependency is load-bearing.

[Inferred] Why it is the right call for their constraint: at sub-second latency over terabytes a day [Search], you cannot run attention over every message. Phishing lures are mostly short, local, and formulaic — "verify your account," "call this number" — which is precisely what an n-gram CNN captures. The inductive bias matches the signal.

“Abnormal combats these tactics by designing our system to run text models on all text extracted from any part of a message—even if the words themselves aren’t inherently malicious”

How does this work concretely and what does it accomplish?

[Post] Motivating example: in callback phishing, attackers "urge victims to call a phone number by sending fake purchase confirmations," varying "the brands they impersonate, the structure of fake receipts, the presentation of phone numbers, or the placement of malicious text," while "the objective remains the same."

[Inferred] What that means for the model: there is no keyword. "Your Norton subscription has renewed for $499.99. To cancel, call 1-800-…" contains no malicious word. Every token is ordinary. The signal lives in:

  • Composition — receipt-shaped document plus unfamiliar sender plus phone number plus urgency, co-occurring
  • Contrast with baseline — this vendor has never emailed this person; this tenant has no relationship with this brand. [Search] VendorBase is the mechanism: "a global, federated database that tracks the reputation of every vendor across all Abnormal customers"

[Inferred] What it accomplishes, precisely: it moves detection from lexical matching (does the text contain bad words) to distributional anomaly (is this text unusual for this sender-recipient pair, and structurally consistent with a known attack objective). That is what makes it robust to attackers varying surface details — the varied things are surface, the invariant thing is the objective, and the objective shows up in composition rather than vocabulary.

“Additionally, we automatically retrain our text models with every message processed.”

What cadence and why? How to avoid catastrophic forgetting?

[Post] As written this implies per-message online weight updates. That is almost certainly not what happens, and the same author says so elsewhere.

[Search] Shiebler, on cadence: "We retrain on a weekly basis because we're trying to take advantage more of customer shifts than attacker shifts." He also draws the distinction the post blurs: real-time updates apply to features, not weights — "Real-time updates being applied to features" vs. "continuous training as referring to real-time updates being applied to the weights."

Accurate reading of the post's sentence: every message continuously contributes to the training corpus; the weights update on a weekly batch cadence.

[Search] Why weekly and not faster — his own reason: the dominant non-stationarity they track is customer-side drift (new tenants, changing vendors, shifting communication norms), which moves slower than attacker behavior.

[Search] And they have a separate fast path for attacker shifts — the genuinely interesting part. On a miss they "Modify the decision layer to adapt to this misclassification. The modified decision layer generates substantial training data," then "Retrain the core machine learning models with this new data." The reason for the split is explicit: "adding new [features] involves retraining the model, which takes weeks."

Two-speed adaptation: decision layer in hours, core models in weeks. The most substantive thing they have published on feedback loops.

On catastrophic forgetting — [Inferred], addressed in neither post.

Catastrophic forgetting is when a model trained sequentially on new data overwrites what it learned earlier. Mitigations, MECE by mechanism:

  1. Don't train sequentially — retrain from scratch over a growing corpus. Weekly batch retraining over a window that includes historical attacks makes forgetting a non-issue by construction. This is almost certainly what they do, and the main reason weekly batch beats online SGD here.
  2. Replay. If warm-starting, mix a reservoir sample of old data into each batch. Cheap, effective.
  3. Regularize toward old weights. Elastic Weight Consolidation (Kirkpatrick et al., 2017) penalizes movement in parameters that mattered before. More relevant to sequential fine-tuning than batch retraining.
  4. Freeze and add. Keep the encoder fixed; adapt only a head or adapter layers. Bounded damage.

[Inferred] The security-specific wrinkle: here you want partial forgetting — dead campaign infrastructure should stop influencing the model. The real requirement is not "never forget," it is "never forget attack classes, do forget attack instances." The operational answer is a permanent regression suite of historical attacks, which their backtesting gives them ("what percentage of historical attacks we would catch today"). You do not prevent forgetting, you detect it, every retrain.

[Inferred] Second wrinkle — training on your own outputs. If model-labeled messages feed the next training set, errors compound. Guardrails: keep confirmed labels distinct from model-inferred ones, weight them differently, and hold out a human-labeled evaluation set that never receives model-generated labels.

“Inbound messages, user-reported submissions, and Detection 360 customer reports all contribute to updates in the neural network models.”

What is the difference between these and are they weighed equally?

[Post] Lists all three, says they contribute, distinguishes nothing and says nothing about weighting.

They differ on the two axes that matter — who produced the label, and what selection process generated the sample:

Source

Who labels

Volume

Label quality

Selection bias

Inbound messages

The system itself (model score / rules), or unlabeled

Enormous — all mail

Weakest; self-generated

Reflects the current model's own decisions

User-reported submissions

End users clicking "report phishing"

Moderate

Noisy both ways — users report newsletters as phishing, and miss real BEC

Reflects what looks suspicious to a non-expert

Detection 360 reports

Customer security teams, then Abnormal analysts

Small

Highest — expert-adjudicated

Explicitly the system's errors — misses and false positives

[Search] What Detection 360 actually is: security teams "submit false positives or missed attacks, and get real-time updates from Abnormal on investigation, conclusion, and remediation," and the fix is global — Abnormal "improves the detection engine to ensure that this scenario does not occur again, for any of our customers."

[Inferred] Weighted equally? Almost certainly not, and they should not be — for two independent reasons:

  1. Label reliability differs by orders of magnitude. An adjudicated Detection 360 case is a near-gold label. An inbound message's label is the model's own opinion — training on that at full weight is self-reinforcement, teaching the model to agree with itself. Standard handling: weight by label confidence, or use inbound mail mainly as unlabeled data for representation learning and reserve supervised weight for adjudicated labels.
  2. Information content differs even more. Detection 360 cases are, by construction, sampled from the region where the model is wrong — the decision boundary. In active-learning terms, maximally informative. There are very few of them relative to billions of inbound messages, so weighting by volume makes them vanish entirely. Upweighting them is what actually closes the loop.

[Inferred] Counter-consideration to have ready: aggressively upweighting error cases shifts your training distribution away from the deployment distribution, which can hurt calibration and inflate false positives on ordinary mail. The clean framing is importance weighting — you are deliberately training on a shifted distribution, so either correct for it or at minimum evaluate on the true distribution rather than the training one.

[Search] A class-imbalance forcing function: they state "Less than 1 in 10 million emails is advanced BEC or lateral spear phishing." At that base rate, uniform sampling of inbound mail yields essentially no positives. Positive oversampling is not optional.

“One powerful component is our rule engine”

How does this fit into the broader system?

  1. The text models produce evidence; the rule engine consumes it. The CNNs in the NLP post output scores over extracted text. Those become attributes. The DSL rule's body_text_contents contains "urgent_language" is a text-model output surfaced as a rule-addressable attribute. The NLP post describes how a rule's text predicates get their values.
  2. They cover different failure modes. Rules are precise, instantly deployable, and brittle — known indicators only. Text models generalize but retrain on a weekly-to-multi-week cadence. Rules cover the window between "new campaign observed" and "model retrained." That is the defense-in-depth claim made concrete.
  3. Both sit under one decision layer. [Search] "Ensemble of Models-as-Features" and "Ensemble of Classifiers," with "a combination of all the above approaches." Text scores, rule hits, and behavioral signals meet in the same feature DAG.
  4. The same FP budget binds both. [Search] "The false-positive rate needs to be as low as one in a million." A rule and a model compete for the same budget — which is why neither ships without backtesting.

How would you improve this system if you were building it today?

1. Encoder — cascade, do not replace. The CNN rationale was speed, and that rationale has weakened: distilled encoders and efficient attention narrowed the gap, and per-token inference cost has fallen a lot since mid-2024. Run CNN on every message (keeping universal coverage) and a small distilled transformer on the top few percent by CNN score or recipient risk. Keeps the cost profile, recovers long-range dependency modeling where it matters.

2. Tokenization robustness — keep the ensemble, add normalization in front of it. Three tokenizers is a good adversarial property but 3× encoder cost. Much of what it buys is cheaper upstream: Unicode NFKC normalization, homoglyph folding (Cyrillic а → Latin a), zero-width character stripping, leetspeak folding. Keep the character-level model as the residual defense for what normalization misses. Also worth evaluating: a byte-level model (ByT5-style or byte-level BPE) as one model with most of the character-level robustness.

3. Representation — put the baseline comparison inside the model, not just downstream. [Post] The real signal is deviation from normal for this sender-recipient pair. Today that comparison appears to happen downstream where text scores meet behavioral features. Push it into the encoder: condition on the relationship (has this sender written to this recipient; what did prior mail from them look like) so the text representation is relative rather than absolute. Retrieval over per-tenant sender history is the concrete mechanism.

4. Extraction — where attackers actually moved. The post mentions images and attachments. Since mid-2024 the growth has been in QR codes ("quishing"), attacks living entirely in the linked page rather than the message, and cross-channel impersonation. Invest here over architecture: QR decoding, headless rendering of destination pages, screenshot analysis with a small vision-language model on suspicious mail. Extraction coverage beats encoder quality when the attacker's move is to put the payload where you are not looking.

5. LLMs — targeted, not universal. [Search] Running LLMs over all mail is "Cost prohibitive," but they are clearly affordable on a filtered slice. Two highest-return uses:

  • Intent classification on ambiguous mail — "is this trying to get the recipient to call a number, move money, or hand over credentials?" That is the objective-level question the callback-phishing example is really about, and exactly what an LLM does well and an n-gram CNN does not.
  • Explanation generation — [Post 1] flags interpretability as the reason rules exist. An LLM writing the analyst-facing rationale for a flag closes that gap without putting the LLM in the verdict path.

[Search] Their own fine-tuned 7B open-source models are the cost-efficient vehicle they have already built for this.

6. Evaluation — the change to argue hardest for. Make adversarial evaluation a standing release gate, not a research project: maintain a perturbation suite (misspellings, homoglyphs, image-embedding, whitespace injection, paraphrase), apply it to known attacks, and require every release to hold recall under it. The multi-tokenizer defense is currently a claim about robustness — without this suite, nothing confirms it still holds after each weekly retrain. Pairs naturally with the attack-generation idea from Post 1.

7. Feedback loop — separate label provenance. Per Q17: keep model-generated labels, user reports, and expert-adjudicated Detection 360 cases in distinct streams with distinct weights, and maintain a human-labeled evaluation set that never receives model-generated labels. Without that, weekly retraining on your own outputs drifts — and you will not see it, because your eval set drifts with it.

2024 Blogpost original

Optimizing search relevance at Instacart using hybrid retrieval

Independent retrieval is suboptimal

My notes

Summary

Context

Recall sets for search are currently constructed by independently fetching documents from lexical search and semantic search, and then independently merging the results.

Problem

Independent retrieval is suboptimal

Solution

Hybrid search architecture that leveraged the concept of query entropy to jointly optimize the recall set generation across the different retrieval mechanisms

Some key insights/new things I learned

  1. Using the lexical and semantic retrieval mechanisms independently often leads to a fixed number of items retrieved from each source, regardless of the query or retailer. This is wasteful and can reduce precision, especially when there are limited relevant items for a query.
  2. Original approach
    1. Use SQL queries for lexical search. The top Kt documents are fetched from Postgres based on these scores.
    2. Use ANN built using FAISS. Bi-encoder model based on the Huggingface MiniLM-L3-v2 architecture used to generate query and document embeddings. At runtime, the query embedding is passed to the ANN service. The top Ke relevant documents are then returned, ranked by the dot product scores of the query and document embeddings.
    3. The top K relevant products after merging these two lists are then passed down to the downstream ranking stages.
  3. New approach
    1. Focused on adaptively tuning the recall set size from each retrieval mechanism based on the request context (query, retailer etc)
    2. Query entropy measures the specificity of a query and models the variation or uncertainty in the number of relevant documents for that query
      1. This is measured by looking at historical conversion data for queries, seeing how much variety there was in the products people bought for a given query
      2. For low-entropy (specific) queries, there are likely a small number of relevant results. For higher-entropy (broader) queries, there are a larger number of relevant results.

What I didn’t understand or am unsure about

Question

Answer

“the ratio of relevant items between text and embedding retrieval also varies by entropy”

Is this something they provide data for in the blogpost?

This isn’t shared in the blogpost.

“The recall threshold for each retrieval mechanism is determined using….”

How does this equation lead to different recall sets for the two retrieval mechanisms? Isn’t the query entropy a general quantity that’s independent of the retrieval mechanism?

Although the query entropy is the same, L, M and Q might be adjusted for each retailer and retrieval mechanism.

How are L, M and Q determined in the adaptive recall set sizes equation?

This isn’t explained in the blogpost. Potential approaches could be:

  • L: size of a standard results page
  • M: determined by p99 Latency SLAs. In a hybrid system, the "Reranker" (the expensive AI model that sorts the final list) has a strict time limit (e.g., 50ms). M could be set to the maximum number of items that the Reranker can process within that window without slowing down the app.
  • Q: set at 95th percentile entropy, making the scaling aggressive enough to be useful for the average search, while ignoring extreme statistical noise

Why was latency reduced by switching to the adaptive recall system?

Not explained in the blogpost but likely related to smaller recall set size for low-entropy (specific) queries.

A smaller recall set size means Reranker finishes faster.

What is the best explanation for why the mean converting position improved?

Larger recall set size for better suited retrieval mechanism means a higher number of relevant items returned.

Smaller recall set size for worse suited retrieval mechanism means less noise in ReRanker (and thus the final ranking).

In what sense Reciprocal Rank Fusion is a “hybrid retrieval” approach distinct from the original approach that they were using in the blogpost?

RRF generates two lists and combines them, which seems to be the same as the original approach outlines.

RRF combines the lists in a more sophisticated way. Instead of taking the Top n from each retrieval mechanism, it takes the reciprocal mean of the ranks from each mechanism and orders the results based on this mean.

In the “Convex Combination of scores” approach, how would they compute the global weights for the retrieval mechanisms?

They don’t state it in the blogpost but they could:

  • Bayesian Optimization / Grid Search with different combinations
  • Learning to Rank (LTR) Model: "weights" are coefficients learned by a lightweight model

What would be the best approach for this problem today?

The full pipeline could look something like this:

  1. LLM Query Expansion
  2. Hybrid (Lexical + Matryoshka Embeddings)
  3. Reciprocal Rank Fusion (RRF)
  4. ColBERT / Reranking

We could also use SPLADE or BGE-M3 to replace steps 1-3

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Learn more about term-frequency algorithms/ts_rank,
  2. Learn more about SPLADE and ColBERT
2025 Blogpost original

Merchant Recommendations 2.0 in Afterpay

They wanted to create a single unified framework to improve performance and reduce complexity (which would in turn reduce infrastructure costs and make maintenance/debugging easier). Weekly batch offline inference may also hinder…

My notes

Summary

Context

Afterpay has a Recommended section where users are provided personalized merchant suggestions. The current recommendation approach relied on “separate models, heuristics, and manual rankings across multiple regions (US, UK, AU, NZ), each tailored to specific needs and product categories”. The previous system was limited to weekly batch offline inference.

Problem

They wanted to create a single unified framework to improve performance and reduce complexity (which would in turn reduce infrastructure costs and make maintenance/debugging easier).

Weekly batch offline inference may also hinder performance compared to more dynamic updates.

Some key insights/new things I learned

  1. Absence of session-level impression logging means we can’t keep track of negatives correctly.

What I didn’t understand or am unsure about

Question

Answer

What is “lookalike retrieval”? Is it usually heuristic based or more complex?

It’s used to expand the pool of potential recommendations by identifying merchants that are similar to those a user has already engaged with. Once candidates are identified, they are incorporated into the downstream ranking stage.

It can be heuristic based but can also be based on Embeddings/Two-Towers/GNNs.

Why exactly is their original strategy “difficult to transition to online inference at reasonable costs”?

Running calculations on demand for collaborative filtering, faces serious scalability issues as user and item counts grow.

{Other reasons not as clear}

First, they say “To overcome this, we consolidated 10 disparate models into a single, unified deep learning framework”

Later, they say “we have taken a step back to streamline retrieval with a country-level approach”.

What are they actually using?

For retrieval, they use a country-level retrieval approach.

For ranking, they use a single global ranking model.

How are the features and weights for the weighted scoring method for merchants decided? If this isn’t described in the blogpost, what would be some good ways to decide them?

This isn’t described in the blogpost but could come from business intuition. It could also come from testing different weight combinations against historical data and then using a metric like Recall@K.

They could use Poisson Regression to get feature weights too.

What are the dataset and process used to generate the graph in “Retrieval Metrics (recall@K)“

  1. Look at historical session data. For each user, identify the "Target Merchant" (the one the user actually clicked or bought in that session)
  2. Run the retrieval algorithm to generate a shortlist of K merchants
  3. Check if the the Target Merchant is anywhere in that top K list
  4. Calculate the final Recall@K percentage as the % of sessions where the Target Merchant was in the K-length shortlist

What is the “Recall (Distinct Users)” line in the “Retrieval Metrics (recall@K)“ graph and why is it there?

Preventing power-user bias. {Why is this bad though?}

What is meant by “embedded into a n×d vector space” for the Customer Metadata Features and Merchant Metadata Features?

Feature embeddings are learned during the model training process. Embeddings are needed since the model can’t take in raw categorical feature data as input.

A lookup table is created for each feature, with n being the vocabulary size (unique feature values) and d being the embedding dimension. Random numbers are assigned initially. These are updated as we go through the training process (make predictions, calculate loss, update weights, repeat).

Why are Temporal Context Features “embedded to account for temporal dynamics”?

Making time non-linear can be useful for various reasons:

  1. 23:59 is mathematically far from 00:00 but practically very similar
  2. Friday (Day 5 of the week) is probably more similar to Saturday (Day 6) vs Thursday (Day 4)

What could be some examples of “Time-Windowed Counter Features”?

  1. Number of times the user opened the app in the last 7 days
  2. Number of unique users who clicked this merchant in the last 30 days
  3. How many times has this specific user purchased from this specific merchant in the last 364 days?

What specific changes enabled real-time recommendations?

  1. DCNv2 model is called "live"
  2. Online feature store (Tecton)
  3. Temporal Context Features

What exactly is meant by real-time recommendations?

Recommendations are generated at the time the user opens the app, using the current features.

What advantage do real-time recommendations have over weekly batch offline inference?

If users’ shopping habits change between batch updates, real-time is better. For example, if you interact with a new merchant category, the app can respond in seconds rather than days.

It also enables Temporal Context like day of week, time of day, etc.

Does the model have any capacity to track user journeys? Like someone buying one piece of camping gear should be recommended more camping gear.

Perhaps, if we had a feature like user_purchases_camping_7d

Transformer-based models would be able to do this automatically if they’re provided, for example, a list of the last 50 merchants the user clicked.

“The simplification of our architecture has freed up resources” - in what ways was the architecture simplified?

  1. Retrieval moved from combination of different sources to per-country
  2. Ranking moved from country-specific models to single global model
  3. Codebase shared with other RecSys at the company like Cash App

“....has forced reliance on synthetic negatives for model training. These hard negatives”

Why are we using the words ‘synthetic’ and ‘hard’ interchangeably here? Don’t they mean different things?

In this case, they don’t explicitly state that merchants are drawn from the same category as clicked merchants, but that is industry-standard in which case those synthetic negatives are actually hard negatives.

They are hard negatives because they are very similar but weren’t clicked.

In what specific ways does “assuming unexposed items are true negatives (limit) the model’s ability to fully capture user intent within the conversion funnel”?

Signal Erosion: Noise added to data because of presence of fake negatives (merchants not shown treated as negative). Gradients can get pulled in arbitrary directions so the decision boundary becomes harder to learn.

Feedback Bias: The model becomes biased toward merchants that were already successful in previous versions of the app. It "assumes" that anything unexposed is bad, which prevents it from discovering new or "long-tail" merchants that the user might actually love.

“By conditioning on clicks, the model focuses on merchants that have already demonstrated user interest, enabling it to better predict purchases as a downstream event.”

Why is this true?

The model only trains on merchants the user actually looked at. This ensures that every "negative" in the training set is a verified rejection.

“the inherent sparsity of purchase events could skew predictions toward frequent converters” - why is this an issue?

This causes Exploration Suppression.

How exactly does the combined model with P(click|impression) and P(purchase|click) work?

  1. Retrieval is proxy for P(click|impression): Uses the weighted scores to find the "best" 50 merchants that users in a specific country are generally interested in. This isn’t actually a probability though, just a score.
  1. Ranking is for P(purchase|click): Takes those 50 merchants and uses the DCNv2 model to perform the final, personalized sort based on which of those 50 this specific user is most likely to actually buy from.

“the combined model…hindered live performance”

Are any details of the live performance evaluation included here?

No

“the simpler P(purchase|click) mode…, delivered superior performance across all critical business metrics in an online setting”

The AUC ROC chart shows that the combined model was better for 2 of the 3 scenarios. How do we reconcile these two statements?

The Combined Model was trained on the kind of data that exists in the test set but this synthetic data isn’t very reflective of the real world.

It’s also possible that the business metrics are a bit different from AUC. AUC measures the probability that a model ranks a randomly chosen positive (relevant) item higher than a random negative (irrelevant) item.

How exactly do the combined model and simpler model differ in practice?

The 0.860 AUC seems extremely high for a recommendation system. Is this reasonable?

Why did transformers work well in the previous blogpost but not in this one?

The sequential signal must have been weaker in this use-case. User history was not as informative

What would be the best approach to this problem today?

  1. Replace Weighted Scoring with "Two-Tower" Vector Retrieval
  2. Upgrade Ranking to "Final MLP + LLM Embeddings"
    1. Feed merchant descriptions, user reviews, and category tags through an LLM (like BERT or a small Llama model) to create "semantic features
    2. These embeddings are then included in the Input Layer for DCNv2
  3. GNNs could be used in both

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Learn more about DCNv2.
    1. https://blog.lukesalamone.com/posts/deep-cross-net-v2/
  2. Look into tools mentioned in Model Training and Serving Infrastructure section
2024 Blogpost original

Real-time Merchant Recommendations on Cash App with Deep Learning

Model: XGBoost has limited ability to capture intricate patterns and achieve strong personalization, as well as dependence on manual feature engineering. Infra: Batch offline inference (daily) increased compute and storage costs since only…

My notes

Summary

Context

Recommendation system for discovery of best offers with merchants that users care about. Previously using XGBoost and daily batch offline inference.

Problem

Model:

XGBoost has limited ability to capture intricate patterns and achieve strong personalization, as well as dependence on manual feature engineering.

Infra:

Batch offline inference (daily) increased compute and storage costs since only a fraction of users accessed recommendations on a daily basis.

Also couldn’t use essential contextual information, such as temporal features

Some key insights/new things I learned

  1. Narrowed down feature set to <10 features vs 100s that they previously had; omitted time-windowed aggregation features
  2. Two-tower model (customer and merchant) predicting the likelihood of a positive customer interaction P(click or transaction)

What I didn’t understand or am unsure about

Question

Answer

How should we think about combining clicks and transactions into a single outcome variable? What are the upsides/downsides of doing it this way vs having separate models for both or weighting them differently (transactions more valuable than clicks)?

The combined approach is simple (one training pipeline, one set of hyperparameters, and one inference call, so lower infrastructure costs). However, it may over-index on "clickbaity" merchants.

Separate models would allow use-case flexibility - some pages could rank by click likelihood while others could rank by transaction likelihood. However, we might have too little data for the transaction model to work well.

Weighting (in the loss function) might align better with the primary business metric (revenue). However, tuning the weights requires some manual effort and is somewhat arbitrary.

The gold standard approach is multi-task learning - shared "backbone" to understand the user/merchant and then splits into specific "heads" for different actions. At the end, you combine the outputs of the heads based on the current business goal.

How did they select the features used for the deep learning approach?

They don’t go into detail but it seems they prioritized simplicity and efficiency.

The ideal approach could use

  • Statistical Importance
  • Permutation Importance
  • Automated Feature Selection
  • Feature Ablation Studies

Why did they omit time-windowed aggregation features initially? Why did they add them later?

Reason for omission:

They didn’t have an online feature store that could serve these aggregated counters with low latency. Without that, aggregating data across multiple time windows during real-time inference is computationally expensive.

Complex aggregations also often lead to discrepancies where the model "sees" different data during training than it does during live inference. Batch brain and streaming brain might be slightly different in how they handle computations. Features may not be 100% accurate in streaming/real-time.

Reason for later inclusion:

Online feature store was ready. Windowed features allowed the model to react to a user's current trends and recent engagement history,

Why would feature reduction alleviate training and inference discrepancies?

Features that rely on exact timing like historical aggregates (e.g., "Total clicks in 30 days") are the most prone to discrepancy so removing those specifically would help alleviate the issue.

Moreover, observability is simplified and we can manually inspect features to identify and resolve issues.

Having a lot of features may also mean we have to reimplement the logic in a different language when streaming to ensure adequate computation speed. This could cause issues if there any discrepancies between how the languages handle edge cases, rounding etc.

What is meant by feature hashing for out-of-vocabulary (OOV) tokens?

Hashing allows any new ID—even one the model has never seen—to be instantly assigned to a valid embedding row without waiting for a daily retraining

{This is a little confusing but see chat with Gemini and this blogpost to understand better}

Why would expanding the size of the embedding matrices improve the model’s performance for cold start scenarios?

The size of the embedding matrices is the same thing as the number of bins available in the hash space. Expanding the size reduces the number of hash collisions. Fewer hash collisions means less noise

What benefit does using transformers offer over the original approach?

  1. It can model more complex, dynamic interactions. It creates a "context-aware" representation where the meaning of a feature actually shifts based on the other features present in the sequence. Unlike decision trees (XGBoost) that apply the same static logic to every user, the self-attention mechanism computes a dot-product between different input tokens (e.g., "Time of Day" and "Merchant Category"). For example, if it is 8:00 AM, the model mathematically "attends" more heavily to coffee-related tokens
  2. They can capture Long-Range & Sequential Dependencies. Approaches like XGBoost are "unordered"—they see a list of features but don't natively understand the chronological flow of a user's life. Transformers process the entire sequence of user actions (clicks, spends, views) in parallel. Because they use Positional Encodings, they can distinguish between an action taken 10 minutes ago and one taken 10 days ago without losing the relationship between them. The model can spot patterns like "User usually browses for 3 days before purchasing," a temporal nuance that traditional models often miss unless a human manually engineers a specific feature for it.
  3. Reduced Manual Feature Engineering

What proxy tools for model explainability did they use? What could they have used

Used

Side-by-Side Comparison: It allowed engineers to view model predictions right next to a user's historical interaction data

Potential Additional Tools

  1. SHAP:
  2. LIME: explains a single prediction by slightly changing the input data (e.g., hiding a user's recent click) and seeing how the prediction changes
  3. Integrated Gradients
  4. Attention Visualization: visualize the "Self-Attention" weights, showing exactly which items in a user's history the model is "looking at" when it decides to recommend a specific merchant

Why would the model’s performance degrade over time?

Long-term drift

Data drift i.e. if the inputs (X) change while the "rules" of the prediction (Y) stay the same (Different users,

Concept drift i.e. relationship between the input and the target changes (different preferences

Short-term staleness

It refers to the gap between when a user takes an action and when the model is able to respond to that specific action. The "Lag" Problem: If you click an offer at 12:00 PM, but the model parameters only synchronize every 5 minutes, you will receive "stale" recommendations until 12:05 PM.

What evaluation metrics did they use?

Used

No specific ones mentioned

Potential Options

Ranking Metrics

  1. NDCG (Normalized Discounted Cumulative Gain): Measures if the merchants the user actually clicked on were ranked at the very top of the list
  2. MRR (Mean Reciprocal Rank): Calculates the average of the reciprocal ranks of the first relevant merchant (e.g., if the first merchant you liked was at position 2, the score is 0.5).
  3. Recall@K (e.g., Recall@10): Measures what percentage of the merchants a user eventually transacted with were present in the top 10 recommendations shown by the model.

Classification Metrics

  1. AUC-ROC (Area Under the ROC Curve): Measures the model's ability to distinguish between a merchant a user will click on versus one they won't, regardless of the classification threshold.
  2. Log Loss (Cross-Entropy Loss): This is the most common loss function for training these models. It penalizes the model heavily if it is very confident about a prediction that turns out to be wrong.

Calibration Metrics

  1. Expected Calibration Error (ECE): In a financial app like Cash App, it is important that if the model predicts a 10% probability of a transaction, exactly 10 out of 100 users actually transact. ECE measures how closely the "predicted probabilities" match "real-world frequencies."

Diversity Metrics

  1. Catalog Coverage: Measures what percentage of all available merchants are actually being recommended across the entire user base (to ensure the model isn't just recommending Starbucks to everyone).
  2. Intra-List Diversity: Measures how different the merchants in a single user's list are from one another (e.g., ensuring the list isn't just 10 different coffee shops).

Business Metrics

  1. Infrastructure Cost Savings
  2. Computational Efficiency
  3. Customer Satisfaction and engagement

Is the final model they used the right setup for multiple recommendations? Since it ranks by probability scores, wouldn't some degree of cannibalization/redundancy occur in recommendations if each recommendation is only viewed in isolation instead of as part of a collective set?

The model predicts a score ($P(\text{click or transaction})$) for a specific user-merchant pair. To generate a list of multiple recommendations, the system runs this calculation for a "candidate set" of many merchants and then sorts them by the highest probability scores.

Cannibalization/redundancy might occur since similar merchants (e.g., two different taco shops) will have very similar embedding vectors. This often results in nearly identical high scores for redundant items.

To resolve this, real-world systems typically add a final re-ranking or diversification stage to transform that collective set into a healthy mix. (1) Maximal Marginal Relevance (MMR), (2) Whole-Page Optimization, (3) Determinantal Point Processes (DPP)

Is the model retrained? If yes, how often?

“We’re enhancing our training and serving infrastructure to support incremental model training”

The details aren’t described in the blogpost but you can initialize the model with the weights from the previous batch training and then tune its existing parameters using a stream of incoming information for the incremental training.

To avoid catastrophic forgetting, you can keep a small, representative "replay buffer" of old data. Other techniques include Knowledge Distillation, Weight Regularization (EWC), or Parameter Isolation.

How did they handle the cold start problem for new merchants or new users?

{feature hashing for out-of-vocabulary (OOV) tokens}

What would be the best approach to this problem today?

Core Architecture

  1. GNN-based Retrieval: Instead of just looking at a user's ID, you represent the entire Cash App ecosystem as a graph where nodes are users/merchants and edges are interactions. This allows the model to capture indirect relationships (e.g., "users who shop at Merchant A also love Merchant B") much more effectively than standard embeddings.
  2. LLM-Augmented Features (Semantic IDs): Rather than simple feature hashing for categories, you would use LLMs to generate Multimodal Embeddings for merchants based on their real-world brand identity, descriptions, and reviews. This solves the "cold-start" problem by letting the model recommend a brand-new merchant simply because its brand "vibe" matches a user’s style.

Multi-Objective "Listwise" Re-ranking

To solve the cannibalization problem you identified, you would replace the "highest probability wins" ranking with a Multi-Objective Optimization (MOO) layer.

  1. Pairwise/Listwise Loss: Instead of scoring one merchant at a time, you train the model to compare pairs or lists. This forces the model to learn that showing three coffee shops is worse than showing one coffee shop, one grocery store, and one online retailer.
  2. Diverse Utility: You would explicitly optimize for multiple targets at once—engagement (clicks), conversion (spend), and discovery (new category exploration)—ensuring the "collective set" of recommendations is healthy and balanced.

Level of understanding: 😐

To-Do’s/Follow-ups

  1. Understand transformers better
  2. Understand feature hashing better
  3. Understand multi-task learning better
  4. Think about multimodal approach

Additional References

  1. Hashing in Modern Recommender Systems: A Primer
  2. Merchant Recommendations 2.0 in Afterpay
2022 Blogpost original

Introducing Natural Language Search for Podcast Episodes

Improve podcast episode retrieval

My notes

Summary

Context

Search at Spotify used to rely mostly on term matching between query and metadata; to compensate for inexactness of user queries, they used fuzzy matching, normalization, and manual aliases.

Problem

Improve podcast episode retrieval

Constraints

Has to be fast

Solution

Dense Retrieval

  • For queries, use the query text as input to the model,
  • For episodes, use a concatenation of textual metadata fields of the episode such as its title, description, its parent podcast show’s title and description etc

Some key insights/new things I learned

  1. They didn’t want to use BERT because its sentence representations are not very good and it’s pre-trained on English text only. Instead, they used the Universal Sentence Encoder CMLM model.
  2. Fine-tuned using the following data
    1. Create (query, episode) pairs from successful podcast searches from historical queries served by Elasticsearch
    2. Successful attempts at search queries after an initial search came up unsuccessful, creating (query_prior_to_successful_reformulation, episode) pairs
    3. Generate synthetic queries from popular episode titles and descriptions to create (synthetic_query, episode) pairs
    4. Small curated set of “semantic” queries was manually written for popular episodes
    5. Negative pairs generated by using in-batch negatives
  3. They used Hard Negative Mining to improve performance
  4. Recall@1 and Mean Reciprocal Rank for evaluation; Recall@30 and MRR@30
  5. Episode vectors are pre-computed for a large set of episodes using the episode encoder in an offline pipeline
  6. They re-rank the top episodes retrieved by ANN using other features like episode popularity in a final-stage reranking model
  7. They use a vector cache is also used to avoid computing the same query vectors too often
  8. Launched an A/B test that resulted in a significant increase in podcast engagement

What I didn’t understand or am unsure about

Question

Answer

For data used in 2b, how do they know the query prior to successful reformulation was intended to search for the same episode as the final query?

This is an assumption. It could be verified by strict time constraints between consecutive queries or lexical overlap.

How does the Universal Sentence Encoder CMLM work?

{Study separately}

What other methods did they try besides the Universal Sentence Encoder CMLM model?

SBERT is mentioned.

Are there any alternatives to concatenating all of the metadata into a single string? Perhaps something that can account for the fact that each of those metadata components represent different characteristics and may have different relationships with the query? Is there a reason why those weren’t considered?

  • Everything might become n times as slow (training, inference etc) if you split them independently.
  • If you insert special tokens before each component like [TITLE] and [DESC], Self-Attention might be able to learn the different relationships.
  • The reranker can look at specific fields more closely to compensate

If the semantic search model was trained independently of the other sources and the final reranker, could that mean it’s not optimized for being complementary to the other sources and might be redundant?

  • Perhaps mitigated somewhat by reranker determining optimal combination of sources
  • The query reformulation data used to train the semantic model will help it retrieve different results than the keywords

How would the approach change if we were trying to solve this problem today?

  • Use Large Language Models (LLMs) to generate Synthetic Training Data {how?}
  • SPLADE or ColBERT for retrieval; cross-encoder for reranking
  • Use podcast transcripts as well. Embedding them should be possible now that context windows are much larger. {How to embed though?}

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Understand vanilla BERT, sentence BERT and Universal Sentence Encoder CMLM better
  2. Understand deployment better
  3. Understand SPLADE, ColBERT and cross-encoder better (https://gemini.google.com/app/a6d3197fd1be5e3d?pli=1)
  4. Code

Code/Implementation

Introducing Natural Language Search for Podcast Episodes.ipynb

Additional References

How Spotify Uses Semantic Search for Podcasts

https://www.kaggle.com/datasets/listennotes/all-podcast-episodes-published-in-december-2017

2018 Blogpost original

Listing Embeddings in Search Ranking

Improve Similar Listing Recommendations Real-Time Personalization in Search Ranking

My notes

Summary

Context

Potential guests can explore Airbnb’s listings through search results generated from an ML model that uses 100+ signals to decide how to rank listings on the search page. Once a guest views a home they can continue their search by either returning to the results or by browsing the Similar Listing Carousel, where listing recommendations related to the current listing are shown. Together, Search Ranking and Similar Listings drive 99% of our booking conversions

The existing algorithm for Similar Listings consisted of calling the main Search Ranking model for the same location as the given listing followed by filtering on the same price range and listing type as the given listing.

Problem

  1. Improve Similar Listing Recommendations
  2. Real-Time Personalization in Search Ranking

Solution

Listing Embeddings learned from search sessions that allow us to measure similarities between listings

Similarities between listings is then leveraged in real-time personalization in search ranking by upranking listings similar to those users have viewed in the last 2 weeks and downranking listings similar to those they skipped

Some key insights/new things I learned

  1. Embeddings trained through Negative Sampling
    1. Positive context listings are within a sliding window of length 5 and negative context listings are randomly sampled from outside that window.
    2. Booked listings are included as global context
  2. Cold-start embeddings created by finding 3 geographically closest listings that do have embeddings, and are of same listing type and price range as the new listing, and calculate their mean vector
  3. To understand what characteristics were captured by the embeddings, they
    1. Performed k-means clustering to see if geographical similarity is encoded
    2. Calculated average cosine similarities for same category listing pairs vs different category listing pairs
  4. To evaluate, they
    1. Calculated the average ranking of the booked listing produced by the embeddings (based on highest cosine similarities between most recently clicked listing and the listing candidates that need to be ranked)
  5. For real-time personalization in the search ranking, they
    1. Wanted to show more listings that are similar to the ones they thoughts users liked
    2. Wanted to show fewer listings that are similar to the ones they thoughts users disliked
  6. They defined the set of skipped listings to include any listing for which a lower positioned listing was clicked by the user
  7. To evaluate if the new model learned to use the embedding similarity features as intended, they plotted their partial dependency plots.
  8. Before deploying online, they found that
    1. The new embedding features ranked high in ranking model feature importances and
    2. Offline tests showed an improvement in performance metrics on a hold-out set when the embedding features were added to the model

What I didn’t understand or am unsure about

Question

Answer

How was the length of the sliding window for negative sampling determined?

This isn’t explained in the blogpost but the authors talk about using “offline tests” to select hyperparameters so 5 may have been picked based on its performance in this evaluation.

Why would positive context listings mostly consisting of listings from the same market, while the negative context listings mostly consisting of listings that are not from the same market result in sub-optimal within-market similarities?

No explanation provided in the blogpost but probably related to the “Easy Negative" problem. The model overprioritizes location as a differentiator because comparing the location for two listings with the original sampling methods is typically enough to distinguish positive from negative. The other features get drowned out and so when we’re just looking at within market listings, the embedding similarities aren’t as meaningful.

What is the optimization function in the image in the blogpost iterating over?

That is the objective function for updating the embedding of a single central listing during training. It looks at positive pairs (listings clicked before and after the central listing), regular negative pairs (random negative listings drawn from the entire vocabulary), market negative pairs (random negative listings sampled from the same market as the central listing) and the single booked listing for that session.

The training process itself iterates through search sessions in a sliding window manner. For each position of the window, this specific optimization function is calculated to update the vector of the central listing l.

“going as far back as 17 clicks before the booking to the last click before booking” - what does this mean?

Imagine a user clicks on 17 different listings before finally booking one. For each of these 17 specific moments, they took the listing the user was currently looking at and asked: "Based on the embedding of this current listing, how highly would we rank the listing the user eventually booked?".

“The A/B test showed that embedding-based solution lead to a 21% increase in Similar Listing carousel CTR and 4.9% more guests discovering the listing they ended up booking in the Similar Listing carousel.” Did they evaluate whether the total number of bookings increased/if improvements in the Similar Listing algorithm cannibalized bookings from other channels?

They don’t mention this in the blogpost but it might have been incorporated as a guardrail metric in the A/B test.

Why was the 2 week session length chosen?

Subjective perhaps based on how users typically take to plan trips

“Specifically, we calculate similarity between market-level centroids from Hc and pick the maximum similarity” - what does this mean?

Users may like different kinds of listings for different markets so splitting by market is necessary.

It’s faster to calculate 1 similarity score vs n (where n is the number of total listings in that market that the user clicked on)

What would be the best approaches for a similar problem today?

  1. Instead of learning a single vector for a listing ID (as in the blog post), build two separate Neural Networks - User Tower and Item Tower.
  2. Transformers for sequences
  3. Graph Neural Networks
    1. A GNN knows if two listings are similar if they share neighbors (e.g., both are clicked by similar types of users), even if they never appeared in the same session directly

Level of understanding: 🤓

To-Do’s/Follow-ups

  1. Understand embedding training procedure better (how do they get from random vectors to representative vectors)
2022 Blogpost original

Building Airbnb Categories with ML and Human-in-the-Loop

Categorize listings into a set of pre-defined categories. Quality ML model Cover Image ML model.

My notes

Part 2

Summary

Context

https://www.shaped.ai/blog/why-airbnb-made-such-a-big-deal-about-categories

Problem

  1. Categorize listings into a set of pre-defined categories.
  2. Quality ML model
  3. Cover Image ML model.

Constraints

  1. Can’t produce an accurate ML categorization model without a training set of human-generated labels

Solution

Listing categorization: XGBoost binary per category classification models

Cover Photo Selection: Vision Transformer model fine-tuned on human review data

Quality: binary one-vs-all models

Some key insights/new things I learned

  1. Limited human resources for labeling so used weighted sum of indicators to determine good candidates for human review
  2. Candidate expansion was done by obtaining embedding neighbors of top candidates from the rule-based approach
  3. Resampling of holdout set to mimic true listing distribution will get you a much more accurate assessment of your model performance
  4. Interesting and creative feature generation described in “Listing Understanding Signals”
  5. Initially only agent vetted listings were sent to production but eventually when the ML model was good enough, listings with the highest scores were sent too
  6. They trained separate models for each category instead of a single multiclass model; explanation in ML Categorization Model section in Part 2
  7. They used feature dropout to avoid overfitting and limited pattern discovery beyond the rules
  8. By calculating feature importance, they were able to figure out which features (POI in this case) to invest time in improving data quality for; the ML model using improved POI features greatly outperformed the model that used initial POI features
  9. Lake POI refinement was done by inspecting listings that have a high model score but are far from existing Lake POIs and removing wrong POIs by inspecting listings that have a low model score but are very close to an existing Lake POI
  10. The VT model was also used to speed up the human review process by ordering the candidate listing photos in descending order of the VT score

What I didn’t understand or am unsure about

Question

Answer

Who defined the categories and how?

Can a listing belong to multiple categories?

Yes

How was the initial set of indicators for the rule based approach developed?

How were the listing embeddings generated?

{Explained in separate blogpost}

How exactly did they downsample the hold out set to mimic the true listing distribution?

How was the 90% precision threshold for the ML model decided? Why did they prioritize a higher precision over recall?

How does feature dropout work?

What is going on with the easier and easiest negatives?

How were the factors considered in the prioritization of below-threshold listings sent for human review determined?

What is meant by special handling of distance and embedding similarity features not to leak the label?

Why did they use a quality tier for ranking instead of a continuous score?

Was the model retrained once deployed into production?

What is keyword expansion?

{Not explained in blogpost}

Why does high imbalance in category sizes and dominance of category-specific features mean it’s better to train dedicated ML per category models vs a single multiclass model?

Why was the average Top 3 precision on all categories selected as the Cover Image ML model evaluation metric?

What was the basis for deciding that the 70% average Top 3 precision on all categories was satisfactory?

“The best strategy proved to be a combination of those factors” based on what?

What would the best way to approach this problem today?

To-Do’s/Follow-ups

  1. Understand vision transformers better

Additional References

  1. https://www.shaped.ai/blog/why-airbnb-made-such-a-big-deal-about-categories
2021 Blogpost original

Identifying Commute Trips at Lime

My notes

Summary

Goal is to identify commute trips

model the problem as a binary classification - “commute” or “not commute”

While they used many different methods, the 3 that they plot are regularized logistic regression, gradient boosting (LightGBM implementation) and an ensemble model using the first 2.

{to add: why interested in identifying commute trips}

Some key insights/new things I learned

  1. They used ROC over precision-recall curve because the classes aren’t too imbalanced, citing this paper.
  2. The relationship between the features and the target variable may change over time. In this case, mid-covid vs post-covid data may have different relationships. They investigated this by using chronological splits instead of random splits and didn’t observe significant performance deterioration. If the relationship had indeed changed over time, a model trained to predict the relationship in the earlier time period should perform worse in the later time period.
  3. The labeled data didn’t reflect the true future population they would be predicting over since it was collected in a non-random way. They adjusted for this by doing non-random splits of training and test data and modified loss functions.
    1. The sample only consisted of 4 & 5 star rated trips so the non-representativeness of the training data vs the real world could have been a problem. They think this doesn’t matter too much because performance didn’t decline too much when training on 5 stars rated trips and evaluating on 4 star rated trips
    2. Region over/under representation was mitigated by adding weights to the loss function
  4. {to add} some thoughts on active learning and pseudo-labeling

What I didn’t understand or am unsure about

  • Are there any drawbacks to the 5 star/4 star approach they took? What alternatives are there?
  • I wonder what the best practices are for studying whether the relationship between the features and the target has changed over time. Could we instead use a dummy variable for time and see if that feature has (1) a statistically significant coefficient in logistic regression (2) a high SHAP value?
  • How are the weights passed into the loss function to adjust for over/under representation? Is this the best way to do this kind of adjustment?
  • {to add} active learning and pseudo-labeling?

Level of understanding: 🙂

To-Do’s/Follow-ups

  1. Evaluation: ROC vs PR curves
    1. What is the actual math/reasoning behind class imbalance affecting our choice?
    2. Is there a specific threshold for class imbalance that makes PR curves preferable to ROC curves?
  2. Modeling: Ensemble
    1. Learn more about how ensemble models work (why are they better than just using the one model that’s better) and try implementing it in code for a synthetic dataset
  3. Evaluation:
    1. More details for effect of non-representative training distribution on model performance on actual predictions