Work
What I work on
I lead AI at Shakers, a flexible-talent marketplace in Madrid. That is the day job and it is where every opinion on this site comes from. Outside it I speak, sit on advisory work, and occasionally lecture.
If you are trying to get an AI system past the demo, these are the places I am useful.
01
Agent architecture that holds at volume
Where the loop splits, what each step is allowed to reach, and which calls need a frontier model versus a small one behind a fixed contract. Most agents in production are one long prompt with tools attached, and the bill and the debugging both come from that.
02
Eval harnesses, and the measurement theory under them
An eval suite is a measurement instrument, so it has reliability, dimensionality and validity whether or not anyone checks. I came to this from psychometrics, the field that spent a century inferring things it could not observe directly. It is most useful where the ground truth is delayed, contested, or a human decision rather than a label.
03
The stack underneath: routing, tracing, cost per turn
Self-hosted Qwen on AWS via vLLM for privacy-sensitive flows and frontier APIs for the rest. Langfuse and Datadog for traces, Promptfoo for evals. The work is rarely the model. It is knowing which call went where, what it cost, and which seam degraded while nobody was looking.
04
EU AI Act, read as an engineering constraint
Risk classification, the quality management system, technical documentation, and the point where Annex III lands on a data model. I ran the 8-month adaptation programme at Shakers as one of 12 companies in Sandbox IA España, with DG IA and BBVA. It is a build problem before it is a legal one.
05
The org design that decides whether any of it ships
Who owns a capability, what a contract between two teams looks like, and why a roadmap built for features caps what an AI system can become. Usually the real blocker, and rarely the one on the agenda.
In practice
- Talks and panels. Conferences, company offsites, university lectures.
- Architecture and eval reviews. Reading your traces and your eval suite with your team, and saying where the loop and the measurement disagree.
- Regulatory sessions. AI Act readiness with the engineers and the lawyers in the same room, because the answers live in the data model.
- Advisory. Committees and standing advisory work on production AI and evaluation.
Past sessions, with dates and who organized them, are on the speaking page. The research page is where the measurement half of this comes from.
Starting a conversation
Thirty minutes, no deck. Bring the system and the problem you are stuck on. If I am the wrong person for it I will say so on the call and usually know who is not.
Book a 30-min chat→