Europe 2026
58 talks, 8 themes, one fast way through the event — key takeaways, standout quotes, and every recording from AGNTCon + MCPCon.
Overview
Across 58 recorded sessions at AGNTCon + MCPCon Europe, speakers returned to the practical work of taking agents from prototype to production: securing identity and payments, governing autonomy for enterprise compliance, and evaluating and observing agent behavior before and after deployment. Other threads focused on architecture and orchestration for multi-agent systems, context and skill design for usable agent interfaces, and the gateways and transport protocols carrying MCP traffic at scale.</intro> </invoke>
Themes
Keynotes and framing
What a Young Foundation Got Wrong — and Fixed
Worth noting
AIF added a sandbox project stage with a 6-month check-in and 12-month deadline to reach "growth" status or be archived, and opened all working groups to public participation.
Quote
Foundations that catch projects early keep them. Foundations that wait don't.
MCP Challenges & Opportunities
Worth noting
GitHub's new stateless MCP protocol cut basic initialize-request latency from ~14ms to ~6ms versus the old session model, enabling scaling to many more millions of tool calls.
Quote
The scalability isn't a joke, and it happens on every layer.
Photos from this session
Security, identity and trust
Securing the Agentic Universe, One Layer at a Time
Worth noting
Layered monitoring (tool calls, agent runs, fleet-wide) caught a multi-turn prompt-injection takeover of an admin console and a personal assistant's weekly cost spiking from $11 to $1,600.
Quote
Any single agent that is compromised could potentially infect your whole fleet and actually take down your whole fleet.
Photos from this session
Which Controls Still Matter When the Model Stops Refusing
Worth noting
Testing an obliterated LLM in Kubernetes sandboxes, the speaker found SSH-isolated, read-only, network-policy-locked pods blocked exfiltration, while text-stream configuration (as in the AWS Trends CVE) was trivially overwritten by prompts.
Quote
Anyone that relies on guard rails from the model itself is going to fail eventually.
Giving Your Agentic Coding AI a Security Brain
Worth noting
Benchmarks cited show AI-generated code is insecure roughly half the time, and in Snyk's own VulnBench test nearly 50% of LLM-reported vulnerabilities didn't reproduce across repeated scans.
Quote
AI review is a measurement, it's not really a verdict.
Photos from this session
Sandboxing My AI Agent, One Layer at a Time
Worth noting
After layering microVMs, seccomp, copy-on-write filesystems and a DNS-aware firewall around his coding agent, the speaker concluded that decomposing the harness into services and sandboxing only the execution layer scales better.
Quote
Separation of concerns leads to security, by splitting the different pieces of the agent harness, you're able to create a more secure system.
Photos from this session
Exploring WebMCP: What Happens When AI Agents Start Using Websites?
Worth noting
WebMCP lets sites expose structured, discoverable tools so agents stop guessing from the DOM, but it's an experimental Chromium-only API that saw multiple breaking changes in just the weeks before this talk.
Quote
If we will leave it like that, this is a recipe for disaster.
Photos from this session
Beyond the Easy 80%: Bringing Legacy, Spatial, and Locked-Down Data to MCP
Worth noting
Safe Software's FME platform exposes governed workflows as purpose-built MCP tools that anonymize exact addresses to block level, letting Copilot query sensitive geospatial crime data without ever touching raw records or credentials.
Quote
MCP standardized how clients discover and invoke capabilities, but the standard interface does not govern what happens after the request arrives.
Photos from this session
Potential Issues for Cross-Domain Multi-Hop API Calls and Their Solution Proposals
Worth noting
Of five cross-domain token-passing methods for MCP, only grant-chaining (JAG) and MCP authorization-delegation satisfy no-fraudulent-use/no-leak requirements, but still need DPoP sender-constrained tokens and a session-cookie comparison scheme to stop user swapping.
Quote
Sender constrained token can only be used by a client application having the right temporal private signing key.
Photos from this session
A Year of MCP Vulnerabilities: Protocol Gaps and What's Being Fixed
Worth noting
Reproducing real CVEs in Grafana, GitLab, Argo CD, and four AI coding tools, the speaker found a recurring root cause: trust checks that verify a string match instead of a real, issued identity.
Quote
Having the illusion that your server is doing a trust check is more dangerous than not having one.
Photos from this session
What MCP Borrowed From LSP's Packaging — and the Trust Lesson It Missed
Worth noting
He live-demos packaging an MCP server as a signed, SBOM- and SALSA-provenance-attested OCI artifact via MCPB and KitOps, showing a single edited line changes the digest and invalidates the prior cosign signature.
Quote
Packaging is an essential part of the security model, and packaging also needs to come with a distribution protocol.
Photos from this session
Attribution by Design: Skills, MCP, and Where Provenance Gets Built In
Worth noting
Proposes borrowing C2PA content-credential signing for MCP skills, plus a new "interceptors" extension (validator/mutator hooks on a standalone server) to audit and attest skill and tool provenance.
Quote
You can't really guarantee that it matches what the author wrote and signed as part of that skill.
Photos from this session
ID JAG: Solving OAuth Sprawl for Enterprise AI Agents
Worth noting
IDJAG lets an enterprise IDP issue a short-lived token (via OAuth token exchange, RFC 8693) that an MCP server exchanges for a normal access token (JWT bearer grant, RFC 7523), eliminating per-app OAuth consent screens.
Quote
OAuth is the worst form of authorization except for everything else.
Photos from this session
Economies of Scale for MCP and Agents: Why You Need an Identity Broker
Worth noting
Zalando open-sourced (MIT, under its incubator org) an identity broker that keeps both agent and end-user identities distinct, using RFC 8693 token exchange and OPA policies so agents never see upstream provider tokens.
Quote
An agent without an identity is simply nothing else than an anonymous process that holds all your credentials.
Photos from this session
Agents Can Pay, Can They Prove It?
Worth noting
Credit Agent lets Claude, ChatGPT, and Goose pull ISO mDoc/SD-JWT digital credentials from on-device wallets via MCP apps to verify age, membership, and payment, with AP2's intent/cart/payment mandate trio enabling autonomous "human-not-present" purchases.
Quote
We can do incredible things with agents, but when it comes to actually doing something consequential, agents fall short because we don't trust that layer.
Photos from this session
Your Agent Has a Wallet. Who Has the Receipts?
Worth noting
The Accord Project extends the x402 HTTP payment protocol with cryptographically signed legal agreements, embedding obligation IDs in blockchain transaction memos to create an evidence trail linking agent payments to contract terms.
Quote
This entire agentic space is missing the legal layer. It's the wild west at the moment.
Governance, control and compliance
Governed Agent Autonomy: Building a Control Plane for Agentic Systems
Worth noting
She reverse-engineered the leaked Claude Code source to extract five governance patterns (planning, permissions, tool trust, verification, runtime accountability) and reimplemented them as an open-source Rust agent runtime engine.
Quote
I open-sourced the governance patterns that I found within Claude code.
Photos from this session
What Networking Got Right That Agentic AI Risks Getting Wrong: The Case for an Agent Control Plane
Worth noting
Drawing on BGP autonomous systems, the speaker proposes an "Agent Control Domain" abstraction with three artifacts (authority, delegation state, data constraints) and five cross-domain invariants to govern agent crossings.
Quote
Federating agents is not enough. What we need to do is also to federate domains.
Photos from this session
Breaking the Governance Bottleneck: Why Scaling Control Gets Agents to Production
Worth noting
WSO2's new Agent Manager control plane separates guardrails, agent identity, observability and lifecycle management from agent logic, letting teams swap LLMs between dev and prod without touching agent code.
Quote
The battle is not at the models anymore, we're up at the harness and potentially up at the control plane.
Photos from this session
Building a Sovereign AI Governance Stack with Open Source
Worth noting
Since January, OpenAI, Google, and Anthropic alone shipped 32 frontier model versions, making a DIY five-layer governance stack (gateway, registry, access, tracing, guardrails) a full-time maintenance job for most companies.
Quote
The gateway and the registry are the most important. Everything else attaches to that.
Photos from this session
From Advisory to Autonomous: A Staged Model for Agent Adoption
Worth noting
Skipping advisory and semi-autonomous stages to ship directly to enhanced autonomy cost an agent-ordering deployment three months of operator trust despite a working solution built in just three weeks.
Quote
Remember, it took me 3 weeks to build a solution and 3 months to gain that trust.
Photos from this session
Scaling Agents
Worth noting
Booking.com shifted from counting adoption to costing the work of turning a ticket into a merge request, then built vendor-agnostic "profiles" bundling skills and MCP configs with gaps for team-level customization.
Quote
Adoption was about counting people, and consumption is about how much they use and what they do with agents.
Photos from this session
Sponsored Session: Agents, Infrastructure, and the Future of AI-native Applications
Worth noting
Vultr pitches itself as an open, sovereign hyperscaler alternative to AWS—33 data center regions, CPU-run MCP servers with edge inference for EU AI Act compliance, and 50-90% lower cost with 20-33% better performance than AWS/GCP.
Quote
We are not AWS. We are a new hyperscaler purpose-built to give you full control of your data and your open stack.
Photos from this session
Outcome Engineering: Why Your Agentic Architecture Doesn't Matter Yet
Worth noting
Citing Gartner's forecast that 40%+ of agentic AI projects will be canceled by 2027, the speaker argues teams should engineer backwards from business outcomes (margin, revenue, retention) rather than starting from tools and architecture.
Quote
Technical quality can make something look very impressive, but it cannot make it important.
Photos from this session
Legal Implications Under EU Law When Deploying AI Agents
Worth noting
Using three AI agent case studies, the speaker shows how post-deployment behavior changes (e.g. voice emotion recognition) can trigger high-risk classification and shift legal responsibility from provider to deployer under AI Act Articles 12, 25 and 43.
Quote
Every action has a potential legal implication under the AI act but also under the other laws that we've talked about.
Photos from this session
CHAP — Open Protocol for Auditable Human-Agent Collaboration
Worth noting
CHAP defines a shared, open protocol — five nouns, seven core verbs, and eleven optional profiles — that records every human-agent decision in an appendable evidence log, exposed natively as MCP tools or A2A skills.
Quote
Chap does not claim compliance on your behalf.
Photos from this session
Verify, Abstain, or Amplify: A Field Guide To Confidently-Wrong Agents
Worth noting
He proposes classifying agent blockers into three routes - verify against an oracle, abstain and route to the person who knows the fact, or amplify by asking someone with taste - so work continues asynchronously instead of halting on every question.
Quote
We want the agent to write down the questions somewhere, so it's the agent formulating the text, not the person writing the ticket.
Photos from this session
Pull Requests Are Dead, Long Live Peer Review
Worth noting
By reviewing one-page implementation plans before agents code instead of pull requests after, Overmind tripled issues and commits completed and cut lead time from 3.2 to 1.3 days.
Quote
It's a hell of a lot easier to reject a recipe than a fully finished meal.
Photos from this session
Skills Need SemVer Too
Worth noting
Proposes adding semantic versioning (major/minor/patch) plus version logs to agent skills, alongside the existing well-known-endpoint digest scheme, to enable rollback, security-audit pinning, and per-model variants.
Quote
Skills are basically prompt templates that you can ship to your agents that will not bloat your context.
Photos from this session
Evaluation, testing and observability
From Vibes To Data: Evaluating Agents on Your Real Work
Worth noting
Datadog's internal eval platform ("aDEEP") showed Sonnet often outperforms Opus on real tasks, guiding a default-model switch that saved roughly $650,000 a month.
Quote
For a lot of engineering tasks, I would say the models are actually fine. Your context is probably not.
Photos from this session
Testing Agents and Their Tools: Offline Evaluation, Synthetic Tasks, and A/B Experiments
Worth noting
GitHub's MCP team layers offline tool-selection benchmarks, synthetic Harbor-format agent tasks, and online A/B experiments with a 5% p-value threshold, having once shipped a broken tool-prefix cache that spiked costly uncached tokens.
Quote
We unit test fragments of our agentic loop, but not everything can be unit tested.
Photos from this session
Why Agent Evals Miss What Production Catches
Worth noting
Logfire's new 'optimize' feature analyzes failed agent traces to auto-propose fixes to prompts and tool/parameter descriptions, applied live via managed variables without redeploying code.
Quote
We believe that you should ship fast and then get your improvements by understanding what went wrong with the runs.
Distributed Mess: A Production Guide To Multi-Agent Failures
Worth noting
Proposes a four-pillar governance framework (operational health, compliance/cost, quality) plus handoff metrics like context utilization and coverage to catch semantic failures that structural tracing tools like MLflow can't detect.
Quote
The first thing we need to know is actually how the agent is behaving in the first place.
Photos from this session
MAS-Lab: An Open Framework for Spec-Driven, Interoperable Multi-Agent Systems
Worth noting
MAS-Lab is a spec-driven multi-agent framework with a TLA+-verified state-machine runtime that runs the same YAML spec locally or in production, enabling 500-run repeated benchmarks for reproducible research.
Quote
It's not failures that show up in the logs, it's failures that are silent, where the agent thought he had the tool but did not check the result.
Photos from this session
From Opaque to Observable: Tracing Multi-Agent OpenClaw Workflows with OpenTelemetry
Worth noting
Cisco's open-source OpenClaw plugin reconstructs hidden multi-agent telemetry—context assembly, delegation routing, memory, and checkpoint history—into unified OpenTelemetry traces, adding six metric groups for alerting.
Quote
We believe that having no signal is worse than having signals that work most of the time.
Photos from this session
Conformance Testing: Turning MCP's Musts and Shoulds into Working Implementations
Worth noting
The latest MCP spec contains 428 "musts" and 246 "shoulds"; conformance tests are now required for every Spec Enhancement Protocol to be finalized, and 45 contributors helped reach near-100% tier-one SDK conformance by the July release.
Quote
Conformance is essentially a testing framework for us to be able to put SDKs, treat them as a black box, and try to assert behaviors about how they handle the spec.
Photos from this session
Coding Agents That Improve Their Own Harness
Worth noting
The speaker describes using a meta-agent to generate multiple harness variants per task, evaluated externally and tracked via cost-per-accepted-change to catch diminishing returns before credit-card costs spiral.
Quote
If that cost is rising and your progress is stalling, then you're hitting a point of diminishing returns.
Fighting Code Slop: The State of Software Factories
Worth noting
On SlopCodeBench, which tests iterative feature-building rather than one-shot tasks, the best model (GPT-5.5) scored only 14.8%, showing models haven't gotten meaningfully better at architecture or maintainability.
Quote
Unattended models will not improve or maintain your code base quality over time.
I Was the Bottleneck, Not the Agent
Worth noting
Running eight coding agents in parallel just shifted the bottleneck to him, so he redefined "done" as explicit, provable claims verified by an agent-run deploy-observe-diagnose loop with logs and traces.
Quote
Prove the behavior or tell me what's stopping you.
Photos from this session
Agentic AI for Enterprise Mainframes: From Dead Code Elimination To Business Knowledge
Worth noting
By pairing MCP-connected static analysis with a human-approval gate, the speaker's dead-code agent removed 3,000 unused lines across seven COBOL modules with zero production incidents, and a lineage agent traced 134 fields across 33 tables with a 10/10 evaluator score.
Quote
The main thing is we have to measure everything, because if you can't measure any output, you can't trust it.
Photos from this session
Agent architecture and orchestration
The Modern AI Stack: Agents, MCP and Skills
Worth noting
Argues that downloading pre-built skills from the internet is now largely unnecessary as models have improved enough to handle most tasks directly, and warns it's also a security risk.
Quote
MCP is now a global standard, it's not just Anthropic's baby, it belongs to all of us to a certain degree.
Photos from this session
Everything Wrong with AI Agents Running Your Business
Worth noting
An agent testing a live API dispensed 30 real drinks and later ordered cup noodles that didn't physically fit the slot, showing that "task completed green" doesn't mean it actually worked in the physical world.
Quote
The moment AI is running a business, what you see is that it's not my agent, it's ours.
Agents Talking To Agents: MCP, A2A and the Reality of Multi-Agent Orchestration in Production
Worth noting
The speaker outlines four control-plane boundaries (context transfer, scoped execution authority, outcome verification, recoverable handoff) added after discovering that per-step "success" in A2A/MCP agent workflows masked end-to-end failures in production telco security deployments.
Quote
A workflow will be closed only after telemetry and the orchestrator confirms the intended result.
Photos from this session
A2A Goes Stable: What Changed, Why, and What's Next
Worth noting
A2A 1.0 unifies all protocol specs around a single Protobuf source of truth, adds signed agent cards and multi-tenancy, and keeps 0.3 and 1.0 agents interoperable via SDK-level version negotiation.
Quote
A2A is the protocol built from the ground up for agent to agent communication.
Photos from this session
223 Pull Requests in 11 Days
Worth noting
Using parallel AI coding agents (up to 15 at once, running overnight via git worktrees), he shipped a new open-source Java dashboard's first release in 11 days for about $2,000 in tokens, versus an estimated six-plus months by hand.
Quote
If you go ten times faster than your competitor, it's not ten times better — it's incredibly better.
Building a Chief of Staff Agent at GitLab
Worth noting
Nick built a multi-agent "chief of staff" at GitLab with sub-agents, a reflection-based memory system, and deterministic hooks enforcing tone/fact-check gates, but found delegation only works once agents are fed explicit company objectives.
Quote
Not everything has to be delegated. Sometimes it's faster if you just do it yourself.
Context engineering, skills and agent UX
Most MCP Servers are Empty
Worth noting
Logging per-call agent "query intent" revealed only 23% of calls to a get-tickets tool needed the whole ticket, driving a GraphQL-style field-selection redesign of the interface.
Quote
An agent that has to do a lot of loops to find out what the data means may end up using even more tokens.
Photos from this session
File Systems as the Context Primitive
Worth noting
Bloomberg proposes a new MCP working group adding five Unix-style filesystem operations (list, read, write, search, discover) to resources, citing a drop from ~150,000 to ~2,000 tokens of tool definitions.
Quote
It is to let useful information leave the active context without becoming lost and bring it back when it is truly needed again.
Photos from this session
MCP Doesn't Have a Context Problem
Worth noting
A live demo combining tool search, code mode, and CLI tool calls in the MC Pi harness generated a GitHub contribution animation using only 47,000 of a million available context tokens.
Quote
It still only used 47,000 tokens, which is incredibly small for the amount of data it's actually pulled down.
Photos from this session
We Built an Agent, We Shipped a Compiler. Here's Why.
Worth noting
A 14,000-line, 42-file Go evaluator-critic "compiler" built to make an AI skill deterministic was ultimately scrapped for a 54-line skill using progressive disclosure and a single guardrail query.
Quote
If you build something which has a high barrier to entry for the team, then adoption is going to lag.
Photos from this session
Beyond Chatbots: Agentic UI With Open Standards
Worth noting
The talk walks through three complementary open protocols—AG-UI for agent-frontend streaming, A2UI for dynamic UI generation, and MCP apps for sandboxed tool widgets—noting AG-UI's npm package now peaks near 2 million downloads a week.
Quote
To chat or not to chat, that's not the question, because everything begins with the user intent.
Photos from this session
Two Users, One App: Designing MCP Apps for Humans and Agents
Worth noting
MCP apps now have two coexisting users (human and agent), so developers must explicitly call "update model context" to sync front-end UI state changes back to the agent, since this isn't automatic.
Quote
There is a function that allows you to give the model context about what you did in the front end.
Photos from this session
MCP infrastructure: gateways, transport and protocol
What Does It Take To Ship a New MCP Spec
Worth noting
MCP grew from ~12 million to over 3 billion package downloads and now has 2,000+ contributors, driving a formal spec process with interest/working groups, SEPs, and soft/hard freeze release cycles.
Quote
Just shipping an MCP spec itself is not enough. Like it's just words on a web page.
Photos from this session
Stateless: The Future of MCP Transports
Worth noting
MCP's new spec eliminates the initialize handshake and sessions entirely, moving to per-request state and explicit state handles so Google's MCP servers can scale horizontally to trillions of requests.
Quote
You don't have to call initialize anymore, you can start talking to your server directly.
Photos from this session
Call Now, Fetch Later: Durable MCP Tasks on an Event Log
Worth noting
In a prototype airline-rebooking flow, a Redis-backed MCP task store duplicated ticket and hotel-voucher bookings after a crash or elicitation pause, while an event-log (Kafka) backend resumed exactly where it left off.
Quote
An event log is an append-only, ordered, and durable sequence of records, events, with producers and independently positioned consumers.
Photos from this session
From Custom to Open Source: Spotify's MCP Gateway Journey
Worth noting
Spotify's single MCP gateway grew from 50,000 tool calls/month to about 4 million/month with 100% employee adoption, routing ~150 MCP servers, and was rebuilt on Envoy's native open-source MCP filter instead of custom code.
Quote
We started with about 50,000 tool codes per month in December last year, and we're now at about four million per month.
From MCP Playground to Org-Wide Infrastructure: Lessons From Building Booking.com's Agent Foundry
Worth noting
Booking.com's progressive tool disclosure cut context usage for 10 enabled MCPs from 189k tokens (18% of context) to just 538 tokens, while its internal platform grew to over 6,000 users organically.
Quote
Fabric now reaches more than 6,000 users at Booking.com, and I have to tell you this is not enforced.
Photos from this session
Your Agents Need a Router: One Integration for Every Model and Tool
Worth noting
Agent Router (formerly Envoy AI Gateway) decouples credential management, model failover, and token-based quota/cost-attribution from agent code via declarative Kubernetes-style policies on an Envoy data plane.
Quote
One of the objectives of agent router is to help you keep your agents boring.
Photos from this session
Photos
Photos


























































































































































































































































































































































































































































































































































































































































































