0:00
/

How This CTO Lets His AI Agents Redeploy to Production Without Him

When Rose's chatbot gets something wrong, nobody opens a pull request. An agent fixes it, tests it against an evaluation set, and pushes it live - without a human touching the code.

A visitor types a question into the chatbot on Cactus Inbound’s website: can you give me a cocktail recipe? The Rose bot politely explains that it’s a marketing assistant, not a bartender, and steers the conversation back to digital marketing. Cactus Inbound is a fake competitor site Benoit Pothier’s team built from scratch, just to demo Rose without touching a real client’s data.

Benoit doesn’t like the bot’s answer. He flags it as inaccurate. Within seconds, a ticket appears in Linear. Somewhere behind the scenes, an agent picks it up, edits the bot’s knowledge and skills, checks its work against an evaluation dataset, and pushes the fix straight to production. No pull request. No code review. Benoit never opens an editor.

In this episode of Build with AI, Tanguy Goretti, CTO at Hexa, sits down with Benoit Pothier, CTO and co-founder of Rose, to trace exactly how that loop works, from a bad answer flagged by a user to a production deploy, with no engineer in between.

Rose is a suite of AI agents born out of Hexa. Instead of a static chatbot, it holds a conversation with visitors, works out who they are, and hands qualified leads off for nurturing and sales outreach. Benoit and his team describe it as an “agent factory”, because the harder problem isn’t answering one visitor’s question, it’s keeping hundreds of these agents accurate as they run.

The harness: agentic, but on the clock

Rose’s core isn’t a simple RAG lookup, it’s a two-phase agentic harness built around one hard constraint: nobody waits minutes for a website chatbot to answer. The first phase pulls in context, screens for spam, and selects the skills and knowledge the response will need. The second phase generates the actual reply. The whole thing is built to land somewhere between 2 and 6 seconds, depending on how complex the question is and how the underlying model providers are performing that day.

“It’s a real agentic behavior, with a very intelligent system - not just going to fetch a basic RAG.” — Benoit Pothier

Structured knowledge beats loose chunks

Feeding that harness is what Rose calls the context engine. It ingests everything a client hands over (a scraped website, sales decks, even transcripts of calls between a rep and a prospect) material that’s often incomplete or flatly contradictory. Rose curates it, then builds two things from it: a standard vector store of chunks, and a graph of entities and relationships (clients, for instance, each represented as their own node).

The graph is what makes a question like “who are your clients” reliable. Pure chunk retrieval struggles with that kind of query, it’s built to find text that’s semantically close to the question, not to guarantee completeness. Querying the graph directly for the “client” entity type does.

“Much more deterministic than a vector search.” — Benoit Pothier, on why the graph handles these queries better than chunks alone

200 skills, kept on a leash

Rose currently runs about 200 skills (pricing guidance, competitor-comparison behavior, gated-content rules, and more) some shared across clients, some specific to one. An LLM selects which skills apply to a given conversation, but that selection is bounded by a rule engine, not left to free reasoning.

That’s a deliberate choice, and one Benoit pushed back on when Tanguy raised a comment attributed to Claude Code’s creator: that with more capable models like Opus 5, teams should delete their skills and let the model figure things out. Benoit’s answer was that newer models are more capable, but they explore more turns to get there, and Rose has to land an answer in one turn to hit its speed target. So instead of letting the agent reason its way to a plan, the system tells it what to do.

“We’re not leaving the agent to work out how it’s going to do it. We tell it what to do.” — Benoit Pothier

From a flagged answer to a Linear ticket

Back in the demo: Benoit opens Rose’s backoffice, finds the Cactus Inbound conversation where the bot dodged the cocktail question, and flags it as “not accurate,” noting what the bot should have said instead. That single action creates a ticket in Linear automatically.

The fix redeploys itself — after clearing the evals

An agent picks up that ticket, edits the relevant knowledge, skills, or code, and reruns the fix against Rose’s evaluation datasets, some scoped per client, some global, some built around a specific feature like the initial conversation classifier. If the fix clears its evals, it redeploys to production on its own. Rose closes roughly 20 tickets a week this way.

A second loop: catching problems before anyone complains

A separate loop watches the system itself rather than waiting on user complaints. An instrumentation layer (Grafana-style dashboards and alerts) tracks metrics like response latency; if the slowest 10% of requests cross 15 seconds, it fires an alert an agent can investigate, with access to the relevant client’s database and logs to diagnose, and sometimes reproduce, the problem. Not every alert needs a fix: Benoit described one case that traced back to a model provider having a bad day, where the right call was to do nothing. A human still checks the diagnosis before anything changes.

Takeaways

Constrain the agent, don’t just prompt it. Rose’s speed and consistency come from a rule engine that bounds what the LLM can choose, at both the skill-selection and response-generation stages.

Structured retrieval beats bigger context for factual reliability. For “how many clients do you have”-style questions, an entity graph gives a complete, deterministic answer that pure vector search can’t guarantee.

Every automated fix earns its way to production through the same eval gate a human would use. The agent that resolves a ticket doesn’t get to skip the check that would catch a regression.

Evals are what make autonomy safe, not optional. As the system does more on its own, the bar isn’t “trust the agent”, it’s “trust the evals that gate what the agent is allowed to ship.”

The bottom line

The interesting part of Rose isn’t the chatbot a visitor talks to, it’s the layer of agents working underneath it, watching for bad answers and infrastructure anomalies, and closing the loop without waiting for an engineer to pick up the ticket. That only works because every automated change, whether triggered by a user complaint or a latency alert, has to clear an evaluation gate before it ships.

It’s also a quiet rebuttal to a narrative going around as models get more capable: that skills, rules, and evals become training wheels you can take off. Benoit’s read is closer to the opposite: the more autonomy you hand an agent, the more precisely you need to define what “correct” means and verify against it, especially when the constraint isn’t intelligence but speed. Rose’s next step is teaching a supervision layer to optimize its agents toward business outcomes like conversion rate, not just correctness - while trusting the same eval discipline to keep it safe to deploy without a human re-checking every change.

Subscribe to receive future episodes of Build with AI.

Discussion about this video

User's avatar

Ready for more?