“In the case of AI agents, Cresta suggests how to modify the agent’s prompt or other parts of its configuration and then runs tests to validate that the agent’s performance improved as a result.“ Does “Cresta” here refer to a Cresta employee or a Cresta system? In other words, is this done manually or automatically? | [Source] The post doesn't say. The sentence is built as a parallel to the preceding one about human agents, where Cresta the platform delivers hints and coaching, so the grammar points toward the system. [Inference] The two verbs are doing different amounts of work. "Runs tests" is plainly automated — that's the evaluator toolkit and simulation suite described later. "Suggests how to modify" is the softer claim, and most likely means the platform surfaces a diagnosis (this turn failed this behavior, here's why) while a person decides the actual edit. The post's own Lessons Learned section supports that reading, since its entire argument is that humans stay in the loop on targeted decisions. So the honest summary is automated detection and automated validation, with a human-authored change in between. "Suggests" is the kind of verb that lets a vendor describe either arrangement. |
Are Cresta’s agents primarily text-based or voice? | [Source] This post is modality-agnostic. It refers to "chat and call transcripts" and mentions both inbound and outbound conversations. [Search] Voice is clearly the emphasis across the rest of their writing. Their Voice Virtual Agent announcement says voice is "the top used channel at over 70% of interactions" and criticizes competitors for "focusing exclusively on chat," and their latency engineering post is entirely about the voice pipeline — speech recognition, voice activity detection, speech synthesis. [Inference] Both channels are supported, but the engineering problems they choose to write about are voice problems. The "interruptions" metric listed in this post only makes sense for voice at all. I'd read the post as voice-first with chat along for the ride. |
“Cresta fixes this cold start problem by grounding the design in historical data” Doesn’t cold start imply no historical data? | [Inference] You're right that the term is being stretched. Cold start conventionally means having no data about a new user or item. Here the cold start belongs to the agent, not the enterprise: the new AI agent has no interaction history of its own, while the enterprise has millions of transcripts. So what they're describing is a transfer-learning move — bootstrap the new agent from the behavior of the incumbent human agents. That's a reasonable idea, but calling it "fixing cold start" is marketing language; a stricter description would be a warm start from a related distribution. And the distributions genuinely differ: customers ask different things of a bot than of a person, and they behave differently once they know which one they're talking to. |
What would you do if you actually didn’t have any historical data and wanted to launch something like this? | [Inference] Roughly in order of cost, I'd do five things. First, mine the written artifacts that exist even without transcripts — standard operating procedures, help-center articles, canned agent responses, past ticket subject lines, and abandoned queries typed into a chat widget. Those give you intents without needing dialogue. Second, have subject-matter experts record twenty to fifty think-aloud role-plays, which becomes the seed corpus your personas are drawn from. Third, generate synthetic conversations from the procedure documents, but validate them against the role-plays so you aren't only testing the agent against your own assumptions. Fourth, launch deliberately narrow, or run the agent in shadow mode alongside human agents so it produces responses that get scored but never sent. Fifth, instrument heavily, because the first few thousand real conversations are the actual asset and you want the relabeling loop to be cheap from day one. The thing to watch is that the first three all encode the same prior. Only real traffic breaks it. |
“finds the major clusters of customer issues” How? | [Search] Their Insights announcement is more specific: Topic Discovery uses "unsupervised LLM clustering on conversation reasons," and generates "an exhaustive taxonomy and up to three levels of subcategories" without teams having to predefine categories. [Inference] In practice that reads as a two-stage pipeline. An LLM first extracts a short conversation reason from each transcript, compressing a long call into a normalized phrase. Clustering then runs over those compressed reasons rather than over raw transcripts — most likely embeddings plus a clustering algorithm, with an LLM naming clusters and merging them hierarchically into the taxonomy. Clustering on extracted reasons instead of full text is the important design decision, because small talk, greetings, and hold messages would otherwise dominate the similarity signal and you'd end up clustering conversations by style rather than by problem. |
Why do so many of their cluster examples sound very similar to each other? | [Inference] Two explanations, and they have different implications. If the screenshot is showing sibling leaves beneath one parent node, then similarity is expected and fine — "billing, unexpected charge" and "billing, duplicate charge" are supposed to differ by one discriminating detail, and the three-level taxonomy makes that likely. The less comfortable explanation is under-merging, which is a real failure mode of LLM-driven taxonomy induction. The model finds distinctions that are linguistically genuine but operationally meaningless, because nothing in the clustering objective penalizes redundancy. The test I'd apply: if two clusters would route to the same conversation flow and trigger the same tool calls, the taxonomy is over-split for the purpose it's being used for. Your suspicion is worth holding onto. |
“how human agents resolved them” Do we know for sure that the human agents' approaches were correct? How? | [Inference] "Proven" is doing unearned work there. Transcripts record what humans did, not what was correct. There are three ways to get closer to correctness, none of them free. You can filter by outcome rather than by behavior, keeping only conversations with a good resolution signal such as no callback within some window, positive satisfaction score, or no subsequent refund reversal — which turns plain imitation into outcome-weighted imitation. You can stratify by agent and learn only from the top quartile on the relevant metric, which is the standard contact-center move and something Cresta's quality management product already scores. Or you can treat the mined flow as a hypothesis rather than ground truth and put it in front of subject-matter experts before it becomes the spec. There's also a survivorship risk worth naming: shortcuts agents use to hit handle-time targets look indistinguishable from best practice in a transcript. |
“Insights surfaces real customer phrasings and look-alike intents” What does this mean? | [Inference] They're two separate things. Real customer phrasings means the utterances people actually produce, as opposed to how the business names the intent internally — customers say "I got charged twice," while the internal taxonomy says "duplicate transaction dispute," and it's the former you need in the router's examples. Look-alike intents means pairs that sit close together in surface form but need completely different handling. "I want to cancel my order" and "I want to cancel my subscription" share most of their words and then diverge into different flows touching different systems. Surfacing those pairs during scoping tells you in advance where your router will confuse itself, so you can add a disambiguation turn or supply discriminating examples before the confusion shows up in production traffic. |
“we can explicitly define what the AI agent will handle and what falls out of scope” What are realistic examples of things that would fall out of scope? | [Inference] Exclusions tend to fall into four groups. There's regulated or liability-bearing work — debt collection promises, insurance coverage determinations, medical or legal advice, anything that creates a binding commitment on the company's behalf. There's high blast radius — refunds above a threshold, account closures, changing a payment method, anything hard to reverse. There's emotionally loaded contact — bereavement, fraud victims, threats of self-harm, and retention saves where human judgment carries the conversation. And there's the long tail, meaning intents that appear a handful of times a month where the cost of building and maintaining a flow exceeds the containment you'd gain. By raw count the long tail is the largest category. The first two are the ones that get written into the contract. |
How is “automation readiness” determined? | [Source] This post lists it as something Insights quantifies, with no method given. [Search] Their dedicated Automation Discovery post is more forthcoming: "The score is driven by structural and operational signals observed in the data, including the frequency and nature of deviations, as well as the level of inferred integration complexity." [Inference] Unpacking those two signals, deviation frequency is a proxy for how well-structured an intent is — if most conversations follow one dominant path, a state machine can capture it, and if they scatter, the intent needs judgment. Integration complexity is a proxy for build cost, meaning how many backend systems the flow has to touch. So readiness is roughly path regularity divided by integration cost, weighted by volume. Worth noticing that this is a build-feasibility score, not a risk score. Nothing in it speaks to whether automating a given intent is safe, and the post's framing quietly blends the two axes. |
“Automation Discovery reconstructs representative dialogue flows from the transcripts” If this tool is looking at a very large number of transcripts (presumably not fitting in a single context window), how is it generating the desired information correctly? | [Source] Not addressed here. Their Automation Discovery post says only that they use "LLMs together with deterministic analysis." [Inference] Nobody puts a million transcripts into one context, so the shape is almost certainly a four-stage pipeline. First, per-transcript extraction, running one model call per conversation in parallel, compressing each into a structured trace: the reason, an ordered list of steps, tool-like events, and the outcome. This is the only stage that sees raw text and it's parallel. Second, cluster the traces rather than the text. Third, align the step sequences within each cluster, which is closer to process mining than to language modeling and can be done deterministically with frequent-sequence mining or transition counting. Fourth, bring an LLM back in only at the end, to name states and write descriptions over an aggregate that now fits comfortably in context. Their phrase "deterministic analysis" most plausibly points at the third stage. The quality risk sits in the first: whatever schema the extractor uses silently bounds what the final state machine is capable of expressing. |
“State machine skeleton: It identifies the dominant paths customers take to reach resolution” Per intent, or more general? | [Inference] Per intent, almost certainly. A single global state machine spanning all customer issues would be close to useless, because the states for rescheduling an appointment and disputing a charge share nothing except authentication. The pipeline ordering in the post supports this: Insights clusters issues first, then Automation Discovery reconstructs flows, so the natural unit is one flow per cluster, with shared preamble states like greeting and identity verification and shared exit states like escalation factored out as common subgraphs. Their decentralized-architecture post is consistent with this, since it names sub-agents for "authentication, policy verification, troubleshooting workflows, payment handling" — which is exactly a shared-preamble-plus-per-intent-body decomposition. |
Is a human the primary consumer of the output of the Automation Discovery tool to guide the development of the agent? Or is it being directly fed into a system somewhere? | [Search] Both, according to their Automation Discovery post. It supports human inspection, and it "can export a structured draft prompt derived from the workflow, serving as a scaffold for AI agents." [Inference] "Draft" and "scaffold" are the operative words — this is a human-reviewed handoff, not a closed loop. That's the right call given what the artifact is. The mined flow encodes whatever the human agents happened to do, so pushing it straight into production would automate their mistakes at scale. The post you're reading agrees implicitly by calling the scoping output a "v1 spec," which presupposes somebody reviews it. |
“Tool catalog: identifies critical data dependencies and system touchpoints” How can you get this information just from transcripts? | [Search] Their Automation Discovery post answers this directly: "When agents say things like 'let me check your order,' 'I'll submit a claim,' or 'I need to verify eligibility,' those patterns often indicate interaction with backend systems." [Inference] So it's linguistic evidence that a system call happened, not observation of the call itself. Two further signals are usually available and they don't mention either: long silences or explicit hold events in a call mark a lookup, and the shape of the data the agent reads back — an order number, a balance, a ship date — reveals both which system was queried and roughly what it returned. The limitation is real, though. This gives you a hypothesis list of integrations with no guarantee of completeness, no API contract, and no visibility into any system the agent used without narrating it. Screen recordings or CRM audit logs would be the ground truth; transcripts are the cheap proxy for it. |
How is “Tool use and data needs” different from “Tool catalog”? | [Source] The post lists them as separate bullets and the distinction it draws is thin. Tool catalog identifies "critical data dependencies and system touchpoints" and enables function schema specification; tool use and data needs catalogues required internal systems and data such as CRM lookups and order adjustments, and informs function call specifications. [Inference] As written these overlap heavily, which by a MECE standard is a defect in the post rather than a distinction you're missing. The most charitable reading is that the catalog is the inventory — which systems exist and get touched — while tool use and data needs is the contract, meaning what each call requires as arguments and what it returns. Inventory versus signature. |
“that might merit separate specialized agents” Is the appeal of separate agents a clean context window or different capabilities (whether through a different model, harness or skills)? | [Source] This post doesn't say. [Search] Their decentralized-architecture post makes the case, and the reasons it gives are mostly neither of your two options. It cites avoiding a single point of failure, since if a central orchestrator "falters due to an unexpected scenario, a model update, or a spike in call volume, the entire experience may degrade"; avoiding errors that "accumulate across long conversations"; and latency, because "every decision and turn reroutes through the orchestrator." It does also describe keeping "the AI focused on the immediate step," which is your clean-context argument. [Inference] In practice I think prompt focus is the dominant real benefit, so your first option. A sub-agent's prompt contains only the rules for its own task, which raises instruction-following accuracy and cuts token cost. Independent testability is second, and this post names it outright — "each skill can be built and tested independently." Running different models per sub-agent is easy to do and occasionally worth it, a small fast model for authentication and a stronger one for troubleshooting, but it isn't the headline reason. |
“and a world state the agent must navigate: authentication, entitlements, status, partial data, the messy stuff” Why is this in the simulated visitor spec as opposed to the test spec? Does their version of visitor both the simulated user and the task? | [Source] Yes, it bundles them. The simulated visitor spec as described carries four things: a persona with tone and vocabulary, a goal defining "what 'done' means," the world state, and behavior policies covering how the person asks, clarifies, and pushes back. Separately the post says "we know 'what should happen' in each scenario," so the expected outcome travels with the scenario too. [Inference] Your instinct that these are separable is right, and separating them would be the better design. Persona, world state, and goal form a cross-product you'd want to sample along independently, so you could hold a persona fixed while varying account state, or the reverse. Bundling them into a single artifact means every combination is a hand-built object and your coverage is limited to whatever someone thought to write down. The likely reason it's bundled is how they're generated: if each visitor is derived from one real conversation, the persona and the world state come from the same source record and arrive already entangled. |
“Because these simulated visitors are generated from de-identified real conversations” How generated? | [Source] No method given beyond the de-identification. [Inference] The plausible pipeline is to take a real transcript, strip personally identifying information using named-entity redaction plus format-based scrubbing for things like card and account numbers, then have an LLM read the cleaned transcript and emit the structured spec. It would infer the persona from how the customer wrote or spoke, the goal from what they were trying to achieve, the world state from facts established during the conversation — were they authenticated, did they have an active policy, was the order partially shipped — and the behavior policies from observed conduct such as pushing back, wandering off-topic, or misunderstanding a rule. Then you discard the original transcript and keep only the spec, which is what makes the artifact privacy-safe: it's a description of a customer rather than a copy of one. Whether they check the generated spec back against the original for fidelity is unstated, and that's the thing I'd most want to know. |
“simulate full multi-turn conversations between each virtual customer and the AI agent” - How to sandbox to ensure no cross-context contamination?
- Different models or same? Does it matter?
| [Inference] (a) The contamination risks are distinct and each needs its own control. Within a single run, the visitor model must not see the agent's system prompt, tools, or internal state — only its own spec plus the dialogue so far — which means two separate processes with separate contexts, not two personas sharing one context. Across runs, you need no shared memory, fresh conversation identifiers, and a reset of the mock backend between runs, because otherwise run N's writes, a payment recorded or a ticket opened, silently change run N+1's world state. The agent's own tools have to hit mocks or a seeded test database rather than production, and this is the one that catches people out, since a tool that looks read-only often still writes an audit record. Finally the grader shouldn't know which configuration produced a transcript, or you introduce scoring bias precisely when you're diffing prompts against each other. (b) Different models, and yes it matters. Using the same model on both sides is the classic evaluation trap: the visitor and the agent share blind spots and failure modes, so the agent scores better than it deserves and the visitor never generates the confusions a different model — or a real person — would produce. Running a different family for the visitor, and ideally a third for the judge, buys you independence. The cost is that your simulated population now inherits whatever quirks the visitor model has, which is exactly why the pool still needs to be grounded in real transcripts. |
“The result of each simulation run is a scored transcript with precise reasons for any failures and artifacts users can diff across prompts, tools, or models.” Same scoring methodology as online and human agents? | [Source] Partly, and the post is explicit about the shared part. It says "a lot of AI agent requirements map exactly to behaviors and are the same for human and AI agents," and that customers running both "can often share the same behavior definitions between the two." That's about Opera behavior definitions applied in production. [Inference] But simulation scoring has something production scoring structurally cannot have, which is a known expected outcome — the post's "we know 'what should happen' in each scenario." So simulation can score outcome correctness by direct comparison, while production can only score behavior adherence by judgment, because nobody knows the counterfactual for a live call. My read is a two-layer scheme: shared behavior evaluators, the Opera definitions, applied everywhere, plus an outcome-comparison layer that exists only offline. That asymmetry is worth holding onto, because a high simulation pass rate and good production behavior scores are not measuring the same thing and can move independently. |
“Because the simulator can spin up hundreds of variations of a scenario, it’s great for catching edge cases that never surfaced in a low-volume manual test.” How to encourage diversity and exploration to uncover more edge cases? | [Inference] Naive sampling from one model collapses toward the mode, and you end up with hundreds of near-identical polite customers. The things that actually work: factorize and grid rather than sample, meaning enumerate the axes explicitly — authentication state, data completeness, emotional register, language proficiency, interruption behavior — and cover the cross-product instead of asking a model for a hundred varied customers. Seed from the tail of real data by sampling source transcripts inversely to cluster frequency, so rare intents are over-represented in the visitor pool. Generate adversarially by conditioning the visitor on a known failure mode, such as insisting on a policy exception, giving the account number with one digit wrong, or changing goal mid-conversation. Apply novelty pressure by embedding generated scenarios and rejecting any that sit too close to an existing pool member, which is cheap and surprisingly effective. And mine production for what you missed, which is what they describe doing post-launch and is the only genuine source of unknown unknowns. The measurable version of all this is coverage over the state machine's transitions: which edges has the pool never exercised? |
“after go-live, we also keep expanding the pool with new conversation patterns as real users interact with the agent” Automatically or manually? How? | [Source] Unstated. The full sentence continues "so we can continually run regression tests." [Inference] The mining step is almost certainly automatic, since it's the same visitor-generation pipeline from question 18 pointed at production transcripts instead of historical ones. The interesting decision is candidate selection, and the natural signals are conversations where the agent escalated, where the customer repeated themselves or expressed frustration, where an evaluator flagged a behavior violation, and where the conversation reason falls outside the existing taxonomy. [Search] That last one is exactly what their Trends feature detects — their Insights post describes Real-Time Trends surfacing "emerging phrases, unusual spikes, and cross-cutting patterns," which is a natural feed into this. [Inference] Promotion into the regression suite should be manual, or at minimum reviewed. An automatically added test encodes whatever the agent did as the expected behavior unless a human supplies the expected outcome, and the post's own emphasis on knowing "what should happen" implies someone has to. |
“encapsulates everything the agent needs: its prompts, decision logic” What is “decision logic”? | [Source] The post doesn't define it. It lists the term alongside prompts, tool integrations, and guardrails, and elsewhere mentions adjusting a "workflow stage" as a thing you can do to a config. [Search] Their decentralized-architecture post is more concrete, describing "deterministic state management" that keeps "track of exactly where a customer is in a process, triggering the right actions at the right moment," combined with dynamic prompting that keeps "the AI focused on the immediate step." [Inference] So decision logic is the non-LLM part of control flow — the state machine from the scoping phase, rendered as configuration. Which stage the conversation is in, which transitions are legal from here, which tool fires on entering a state, when control hands off to another sub-agent, and when the conversation escalates. The language model handles what gets said within a state; the decision logic constrains which states are reachable at all. That's the hybrid architecture they advertise, and it's also why a config bundle is diffable in a way a single prompt isn't: state transitions are structured data, prompt prose is not. |
“Under the hood, the agent’s reasoning and capabilities are extended by Cresta’s AI Agent Framework which is a declarative framework that lets developers register backend functions” What is “a declarative framework”? | [Inference] Declarative means you describe what something is and the framework works out how to execute it, as opposed to imperative, where you write the control flow yourself. Concretely here, a developer writes an ordinary function and annotates it with a schema — a name, a description, typed parameters, a return type — and the framework handles everything downstream: exposing it to the model as a callable tool, validating the model's arguments against the schema, invoking the function, and feeding the result back into the conversation. The developer never writes the "if the model requested this tool, parse its arguments and dispatch" plumbing. Python decorators are the usual surface syntax for this. The real payoff is that the tool's declaration and its implementation are the same artifact and therefore can't drift apart — which matters, because "schema mismatches" is one of the failure types the post names in its optimization loop. |
“The framework is optimized for ultra-low latency” How? | [Source] The post asserts it without method, tying it only to "preserving a smooth real-time conversation." [Search] Cresta has a whole post on this. Their stated budget is that pauses beyond about 300 milliseconds "can feel unnatural" and anything past roughly 1.5 seconds "can rapidly degrade the experience." Component budgets they cite: audio preprocessing 25 to 50 ms, speech recognition 200 to 300 ms, turn detection at least 600 ms, language model time-to-first-token anywhere from 250 ms to over a second, and speech synthesis first byte 100 to 500 ms. The named techniques are streaming APIs with no DNS lookup in the critical path, reused connections to the model, WebRTC transport which they say "can reduce latency by up to 300ms," guardrail model calls issued concurrently with the main call rather than after it, speculative triggering meaning "starting the LLM call before the user fully stops speaking," hedging meaning "launching multiple LLM calls in parallel and using whichever returns first," and smaller models in the live loop since "reasoning models generally can't be used within the live response loop." [Search] On tool calls specifically, which is what the framework sentence is really about, their sharpest observation is that if you look something up before the model can say anything at all, "the 1 LLM first-token latency becomes 1 LLM latency + 1 LLM first-token latency." Their mitigations are to fire predictable lookups concurrently at the start of the conversation, emit filler or wait messages for calls under about ten seconds, and run anything longer asynchronously. [Inference] That last cluster is the actual content behind "optimized for ultra-low latency." The framework's job is to make a registered function's round trip overlap with generation rather than block it. |
“Cresta’s framework supports this composition, so each skill can be built and tested independently” What does skill mean here? | [Source] The post uses skill more or less interchangeably with specialized sub-agent — the preceding sentence is about "specialized sub-agents for distinct tasks." It also uses "micro-skills → agents" back in the scoping section, describing task decomposition that informs multi-agent architecture. [Search] Their decentralized post names the actual units: "authentication, policy verification, troubleshooting workflows, payment handling, and more." [Inference] So a skill here is a bounded task with its own prompt, its own tools, and its own entry and exit conditions — one coherent chunk of the state machine, packaged so it can be exercised in isolation. Worth noting this is not the Anthropic sense of "skill" as a loadable instruction bundle; it's much closer to a microservice with a prompt attached. A good test for whether something deserves to be its own skill is whether you can write a pass/fail condition for it without referring to the rest of the conversation. |
“then unified in an orchestrated flow” By an orchestrator agent? | [Source] The post says "orchestrated flow" and leaves the mechanism unspecified. [Search] Their architecture post argues explicitly against a central orchestrator agent. It says a central orchestrator "can also create a single point of failure," that errors "accumulate across long conversations," and that centralized designs "introduce extra steps, since every decision and turn reroutes through the orchestrator." Their alternative is "a network of specialized subagents" that "handle their areas of expertise directly and in parallel," coordinated by deterministic state management. [Inference] So "orchestrated" here means orchestrated by configuration, not by a supervising model. The state machine decides which skill is active and when control transfers, and there's no LLM sitting above the sub-agents making a routing decision every turn. That's a substantive architectural claim and a defensible one for voice, since it removes a model call from the critical path and makes handoffs deterministic and independently testable. The tradeoff is rigidity: anything the state machine didn't anticipate has no path through the graph and falls through to the escalation edge. |
“The optimization loop is simple: Tweak the agent config or remote functions, run the test set, and fix whatever shows up in red (bad routes, schema mismatches, guardrail violation etc)” Overfitting concerns? | [Source] Not addressed. The post presents the loop approvingly and notes that every run is archived for comparison against the last green one. [Inference] Yes, and I'd call this the sharpest methodological gap in the post. Fixing whatever shows up red and then rerunning the same test set is gradient descent on the test set with a human as the optimizer. Four specific risks follow. You get test-set memorization, where prompt edits patch individual failing scenarios rather than the underlying behavior, which shows up as a pass rate that climbs while production quality doesn't move. Nothing in the description mentions a held-out split the config is never tuned against, which is the minimum defense. With hundreds of stochastic scenarios and repeated runs you get multiple-comparison drift, where some red-to-green flips are noise and chasing them adds prompt text that constrains nothing real. And prompt bloat is the visible symptom of all of it: each patch appends a rule, the prompt grows, instruction-following degrades, earlier rules get quietly ignored, and the fix for failure N breaks passing case M. What I'd expect in a mature version: a frozen holdout set, freshly generated scenarios each cycle, several seeds per scenario with a significance test on the difference rather than eyeballing red versus green, and a regression gate that reruns previously-passing cases. Their post-launch mining partly addresses this by continuously refreshing the pool, which is the right instinct — but it only works if new scenarios are being added faster than they're being tuned against. |
“If it clears the requirements (pass rate, no criticals, tool-accuracy targets)” - What is pass rate?
- What are tool-accuracy targets and why do we care about them?
| [Inference] (a) Pass rate is the fraction of simulated scenarios where the agent reached the expected outcome, with the post's "we know 'what should happen'" providing the comparison basis. It's a conversation-level metric, which their companion post distinguishes from turn-level checks. Two things need pinning down before a pass rate means anything across versions: what counts as a pass, meaning exact outcome match versus evaluator-judged acceptability, and how the scenario pool is weighted, since a pass rate computed over a pool that over-samples easy intents isn't comparable to one computed over a harder pool. (b) Tool accuracy is whether the agent called the right backend function, at the right moment, with correctly-formed arguments. It decomposes into at least four failure modes worth tracking separately: the wrong tool was selected, the right tool was called at the wrong point in the flow, the arguments were malformed or hallucinated — the "schema mismatches" the post names — or no tool was called when one was required. It earns its own gate for three reasons. It's where the irreversible damage lives, since a clumsy sentence is recoverable and a refund issued to the wrong account is not. It's the one dimension where the agent touches systems of record, so its errors escape the conversation and persist. And unlike conversational quality it's cheaply and objectively checkable, because you know what the correct call was, which makes it a good hard gate rather than a soft score. |
“Once an AI agent is live and serving production traffic, we monitor its performance” Who is “we” here? Cresta or the customer? | [Source] Ambiguous, and the referent shifts within the section. The sentence continues "much like we monitor human agent conversations," and human-agent monitoring is done by the customer's quality team using Cresta's tools. But the following sentences say "we monitor all the metrics we care about tracking," and then describe "making it easy for our customers to align the system," which puts customers in the third person. [Inference] Most likely both, with the weight on the customer. Everything described — Dashboard Builder, Opera, AI Analyst — is a self-serve product the customer operates, and the entire framing of Opera is that customers define their own behaviors. Cresta presumably also watches closely during deployment and ramp. The slippage in "we" is worth noticing because it conceals an operating-model question: is post-launch tuning a service Cresta performs, or a product the customer runs? The post reads as the latter, sold with the former during onboarding. |
“This includes metrics common to AI agents, or specific to certain types of AI agents. For example: Interruptions” What are interruptions? | [Inference] An interruption is the customer speaking while the agent is still speaking — barge-in. It gets tracked because it's a cheap, automatically detectable proxy for several distinct problems at once: the agent is talking too long, the turn-detection threshold is mis-tuned so the agent starts before the customer has finished, the response was wrong and the customer is cutting in to correct it, or the customer is simply frustrated. A rising interruption rate is an early warning that requires nobody to read a transcript. Its weakness is exactly that ambiguity — it won't tell you which of those four is happening, so it's a trigger for investigation rather than a diagnosis. |
“Trends and Anomalies page automatically detects shifts in the distributions of conversation reasons” Is this mechanistically different from Insights? | [Search] Their Insights post draws the distinction directly. Topic Discovery provides "durable structure: a trusted view of what customer conversations are about," while Real-Time Trends provides "the real-time signal: visibility into what is changing right now, including events no taxonomy was built to catch." [Inference] So they're layered rather than alternative, and the mechanical difference is the unit of work. Topic Discovery is a batch clustering job that induces a taxonomy: expensive, run periodically, output is a stable label set. Trends and Anomalies is a streaming statistical job over labels that have already been assigned: cheap, run continuously, output is a change signal on distributions you already defined. The phrase about "events no taxonomy was built to catch" implies a second component that isn't just distribution monitoring — emerging-phrase detection over unlabeled text, which catches the thing that doesn't have a bucket yet. That's the piece that genuinely overlaps with clustering, and is presumably a lighter-weight version of it run over a recent window. |
“Show them a specific conversation turn where the agent did something questionable” How do you know in advance if an agent did something questionable without showing it to a human? | [Source] The post doesn't address the selection problem, which is the load-bearing gap in its Lessons Learned section. Their companion evaluator post doesn't either — it covers how to calibrate evaluators, including a preference for binary over numeric scales because "numerical scales are inherently more subjective than binary classifiers," but not how candidates get chosen for review in the first place. [Inference] This is the classic active-learning problem, and the key realization is that you don't need to know the right answer to select well. You need a cheap signal that correlates with being wrong. The usable ones, roughly in order of yield: evaluator disagreement, where you run two or more judges with different prompts or models and route every disagreement to a human; evaluator uncertainty, where the judge's score sits near its decision boundary or its self-consistency across sampled runs is low; hard rule violations, which are detectable deterministically and always worth surfacing — a promise made, a verification step skipped, a tool called before authentication; outcome signals such as escalations, customer repetition, negative sentiment, or a callback within 24 hours, which need no judgment at all because the conversation tells you it went badly; and novelty, meaning turns far from anything in the labeled set, where the evaluator is least trustworthy. The honest framing is that you can't know a turn is questionable, only that it's uncertain or anomalous. What that buys you is throughput: it converts "read a thousand transcripts" into "answer fifty yes-or-no questions," which is precisely the fatigue problem the post's lesson is about. The failure mode to watch is that uncertainty sampling systematically misses confident-but-wrong turns, so a small random audit sample has to run alongside it. |
“Every clarification from a human should either adjust the agent or its evaluation guidelines going forward” Adjust in what way? Prompt? | [Source] Verbatim, and the sentence continues "so we don't ask the same question twice." The mechanism is not specified. [Inference] The sentence names two destinations, and they take different kinds of edit. Adjusting the agent means changing the config bundle, and the post's own framing says a config is much more than a prompt — so the right edit depends on the failure. A wrong tool call is a schema or tool-description fix, not a prompt fix. A missing empathy line is a prompt fix. An illegal path through the conversation is a state-machine fix. A hard compliance rule shouldn't go in the prompt at all; it belongs in deterministic control flow. Adjusting the evaluation guidelines is the other half and it changes the judge rather than the agent: when a human says "no, that response was actually fine," the agent was right and the evaluator was wrong, so the rubric gets amended. [Search] That second half is what their evaluator post is about — aligning stakeholders on what a behavior like "Agent Assumes Payment" means in a given context, where "each misalignment creates an opportunity to refine the guideline." [Inference] The requirement hiding inside "so we don't ask the same question twice" is that each clarification has to be stored as a durable artifact with provenance — a rubric clause or a config line — rather than applied once and forgotten. Otherwise the same ambiguity resurfaces with the next reviewer. |
“The key is to have a process that rapidly discovers and fixes failures when they occur.” What part of their pipeline actually provides rapid feedback? | [Source] The post answers this in the very next sentence: "That's why the simulated testing and the post-launch analytics are so vital." No turnaround times appear anywhere in the post. [Inference] Discovery and fixing have different bottlenecks and it's worth separating them. Discovery is the genuinely fast part, and it's the post-launch layer that delivers it: Trends and Anomalies flags distribution shifts without anyone asking, the Opera behavior evaluators score every conversation rather than a sample, and AI Analyst answers an ad-hoc question in minutes instead of requiring an analyst to pull and read transcripts. Fixing is where the clock slows. The loop is edit config, run the simulation suite, compare against the last green run, lock the version, push. Simulation is the part that parallelizes, but the post never says how long a full suite takes, and that number is what actually sets cycle time. [Inference] What's conspicuously absent is anything between "detected in production" and "fixed in the next version" — no hot-fix path, no feature flag, no way to disable one misbehaving skill while the real fix is built. Config versioning and rollback are the closest thing mentioned, but rollback reverts everything rather than the one broken flow. For a voice agent carrying live traffic, that gap matters more than the length of a test run. |
“We learned to prioritize requirements by criticality.” How does this prioritization actually manifest? | [Inference] It should show up in four concrete places. In the release gate, where the criteria quoted in Q29 are "pass rate, no criticals, tool-accuracy targets" — "no criticals" is a hard blocker while pass rate is a tunable threshold, so criticality manifests as a different kind of gate, not a heavier weight. In where the behavior is enforced, since critical requirements shouldn't live in prompt text at all: you cannot reach the account-change state without passing through verification, and that's a graph property rather than an instruction. In evaluation scope, where their companion post's turn-level checks run on every conversation for critical behaviors while style preferences get sampled conversation-level scoring. And in the agent's own trade-off instructions, which is what the post means by guiding trade-offs — telling it explicitly that if brevity conflicts with a required disclosure, the disclosure wins. [Inference] The open question the post leaves is who assigns criticality and whether the scale is graded or binary. Binary is usually enough and avoids a long argument about tiers. |
In general, how would you design this system to ensure that callers aren’t able to use the generic LLM capabilities and are limited to its customer support role? What guardrails? | [Inference] Layer them, because no single one holds, and note which layers are constraints versus requests. The strongest guardrail is architectural, and their design already provides it. If the agent is a set of narrow skills gated by a deterministic state machine (Q27), there is no state in the graph whose job is "answer general questions," so an off-role request has nowhere to go. A monolithic prompt-driven agent has no such structural limit. The second is capability limitation: the agent can only do what its registered functions do. If no function issues a refund above a threshold, no amount of persuasion produces one. This is the layer that cannot be talked around. The third is a small, fast input classifier ahead of generation, labelling each incoming turn in-scope, out-of-scope, or adversarial and routing the last two to a fixed response rather than to the model. The fourth is output filtering — the concurrent guardrail model call their latency post describes, running in parallel with the main call so it costs no response time, checking for off-topic content, unauthorized commitments, and data leakage. Compliance disclosures and verification prompts should be templated strings rather than generated text, since there is then nothing to manipulate. [Inference] The framing that matters: an instruction in a prompt is a request, not a constraint. Anything you genuinely cannot allow should be enforced by what the system is able to do. |
How would you improve this system if you were building it today? | [Inference] Five changes, ordered by how much they would move real quality. First, fix the evaluation loop, because everything else is measured through it. Freeze a holdout set the config is never tuned against, generate fresh scenarios each cycle, and run several seeds per scenario with a significance test rather than eyeballing red versus green. Without this the pass rate is a vanity metric (Q28). Second, decompose the simulated visitor. Store persona, world state, goal, and expected outcome as four independent artifacts and sample the cross-product instead of hand-building bundled objects (Q17). This is what makes coverage measurable — you can then report which state-machine transitions the scenario pool has never exercised. Third, build the human-review sampler the post assumes but never describes (Q33): disagreement between two judges, plus uncertainty, plus a small random audit sample to catch the confident-but-wrong turns that uncertainty sampling misses. Fourth, split automation readiness into two scores (Q10) — build feasibility and automation risk — and let the risk score act as a veto rather than averaging into a single number. Fifth, add a surgical rollback path. Being able to disable one skill behind a flag, rather than reverting an entire config version, is what turns "rapidly fixes failures" (Q35) from an aspiration into an operational property. [Inference] The one thing I would keep unchanged is the decentralized state-machine architecture. It is the decision that makes every other part of the system testable. |