Welcome
/
blog
/
LabTechDay #19: agentic AI, from design to production
Articles
23.09.2026
5 min

LabTechDay #19: agentic AI, from design to production

On September 23, 2026, Diametral dedicated its 19th LabTechDay to AI agents. For the first time, our teams in Paris, Lyon, Brussels, São Paulo and Bogotá worked on the same programme on the same day. On the agenda: the architecture choices, governance frameworks and field lessons that separate an agent in production from a promising POC.

An agent is a system that pursues a goal using a model, instructions, tools, data and a control loop. We design, build and operate them for our clients. This day allowed us to consolidate our methods and align all our engines around a single doctrine.

LabTechDay, our expertise lab

LabTechDay is a recurring internal format: a full day where our experts compare their engagement practices, test approaches and produce assets that can be reused with our clients. Each edition is led by one engine; this one was led by the Applied AI & Data Science engine.

This 19th edition ran at Group scale: all our offices, on a shared programme. It unfolded in three stages:

  1. Setting the frame: the shift from predictive AI to generative and then agentic AI, the anatomy of an agent (model, harness, tools, memory, guardrails) and the autonomy scale, from simple generator to autonomous agent.
  2. Designing and building: two parallel workshops, one focused on design, the other on engineering.
  3. Testing against reality: a global plenary with quantified field feedback, successes and limits alike.

We came away with a vocabulary shared by all our consultants, reference architectures and an Agent Solution Canvas already usable in the scoping phase with our clients.

Two streams, every engine: design, then build

A reliable agent depends as much on its design as on its engineering. We therefore mobilised all our engines across two complementary streams, brought together at the end of the day around the same case.

Stream A: Design. Designing an agentic solution.
  • Question: Do you need an agent, and under what conditions?
  • Engines: Strategy, Product Delivery, Governance & Quality, Decision Analytics, Acculturation & Change
  • Format: Scoping a realistic client case, then review by a five-role board
  • Deliverable: A one-page Agent Solution Canvas

Before building, we qualify the need. Agent, workflow, copilot or conventional application: we apply three tests. Can the flowchart be drawn in advance? Does the next step depend on what the previous one found? Who executes, the human or the system? When the path is predictable, we recommend a workflow: cheaper, more testable, more predictable. An agent is justified when the path is discovered along the way.

The teams applied this grid to a realistic case: an agent that produces a first draft of a commercial proposal from a request for proposal, discovery notes and past references. They formalised it in our Agent Solution Canvas: 14 decisions across 4 phases (Frame, Grant, Control, Commit), each with a validation criterion. Autonomy level, blast radius of each tool, trust boundaries, human decision rights, quantified value hypothesis, first production scope. Each canvas was then defended before a five-role board: client sponsor, enterprise architect, security and governance lead, delivery lead and end user.

Stream B: Build. Building an agent from scratch.
  • Question: How do you build a reliable, governable agent?
  • Engines: Applied AI & Data Science, Architecture & Engineering
  • Format: Building an agent from scratch: model, tools, control loop, guardrails
  • Deliverable: A working agent harness

Our technical profiles assembled everything around the model: tool calling, the decide-act-observe loop, memory, permissions, observability and execution budgets. Our architecture principle: a probabilistic core inside a deterministic shell. The model brings judgement; the code brings guarantees.

Both streams converge on one point: a control is designed into the architecture, not into the prompt. “It cannot send external emails because it has no sending tool” is a control. “We ask it not to” is not.

Field lessons: what works in production, and what remains to be mastered

The plenary brought all countries together around two cases deployed with our clients, presented with their metrics and their limits.

Luxury & beauty: a product advisor agent for a major house

We designed and deployed an advisor agent on the fragrance and beauty catalogue of a luxury house. The need: answer a question such as “what is an alternative to my lipstick, and how should I apply it for this brand?”, when the answer is scattered across several systems.

  • Scope: more than 4,000 products, more than 200,000 prices cross-referenced by market, country and size, 6 data sources.
  • Architecture: a ReAct agent, 18 tools exposed through 4 MCP servers, text-to-SQL and RAG, deployed on Kubernetes, traced and measured on every run. Each new use case (customer service, retail store) is declared through configuration, without rebuilding the foundation.
  • Results on a 100-case regression test campaign: 76% product recall, 6.9 s median latency, 100% tool precision (no out-of-scope calls).
  • Guardrails: every answer is traceable back to a reference data point. Refusals (medical advice, brand strategy) are tested and versioned just like expected answers.

We also shared the limit we ran into. On the most open-ended questions, the quality target had not been defined upfront by a business owner. Our recommendation: have the teams who own the use case define what a “good answer” is from the scoping stage. That is what makes even the most subjective questions measurable.

Delivery: the Hybrid Squad, two humans and n agents

Adoption of coding agents has taken off; value creation much less so. According to BCG, 5% of organisations capture most of the value from AI, and the gap is organisational, not technological.

Our answer is the “2H + nA” Hybrid Squad: an accountable Product Owner and Data Architect, agents that execute under a versioned contract, and five human validation gates between an agent and production. Three principles shape it: humans remain accountable, agents are bounded, autonomy is earned through evidence. A live demo illustrated the model. This work will be the subject of a book, The Future of AI-Powered Delivery.

Saying no to agent washing

Not every automation warrants an agent. A monthly sales pipeline report follows a fixed sequence: there, we recommend a workflow, with a model limited to the summary step. Qualifying the need honestly is part of our expertise.

Our approach: Diametral Process applied to agents

An agent that holds up in production draws on all our expertise, from strategy to operations. This is what Diametral Process, our methodology for building AI-native companies, structures:

Design: Agent relevance, quantified value hypothesis, scope, autonomy level

Build: Architecture, tools and permissions, guardrails, evaluations in both directions (acting and refusing)

Scale: Extension through configuration, human decision rights, team adoption

Run: Observability, continuous regression testing, named owner on the client side

Our conviction: autonomy can be delegated, accountability never. We help our clients choose the right level of agency, build it and operate it over time.

Other items

See all
Vue aérienne d'un marais avec de petits cours d'eau sinueux traversant des zones de végétation brune et des berges sableuses.

contact

Is your data ready for AI?

A 30-minute exchange with one of our experts to assess your Data maturity and identify the first actions.

Book a diagnosis