What happens when an AI agent's mistake can cost you the client, not just the demo.
You've proven your AI agent works in a sandbox. Getting the business to trust it in production, with real customers watching, is a different problem entirely.
Matthew Sekac is Head of Data & Analytics at Welocalize, where he has spent the past 18 months building an AI agent framework to handle the project management complexity behind thousands of translation and localisation projects. He joined Welocalize in 2020 and has over a decade of senior sales strategy experience in IP and life sciences translation before moving into data leadership.
Matt walks through exactly how his team proved an AI agent framework could handle over a million manual tasks a year without a mountain of rules-based technical debt. You'll get a concrete method for testing agent reliability against human benchmarks before you ever touch production and a way of thinking about human-in-the-loop verification that scales instead of adding new work.
This one covers ethnographic research into how project managers actually work, shadow testing agents against human output, and the evals-based thinking needed to get leadership buy-in for AI at scale. It's built for data and analytics leaders trying to move AI agents from pilot to production, not for anyone looking for a quick AI hype fix.
Key Takeaways
- Welocalize ran a 30-day shadow test where an AI agent mirrored every human decision on live tasks before a single customer saw it, and that proof, not confidence, is what got leadership to say yes.
- The team found their "quality check" agent was inversely correlated with actual accuracy: the more confident it was that a human wouldn't override it, the more likely a human did. The reason why says a lot about where reasoning agents still fail.
- Nobody asks if a human process is error-free before automating it, but everyone expects zero errors from AI. Matt argues that a mismatch is the real reason AI trust conversations go sideways.
- Before writing a single rule, the team spent weeks just watching project managers work, uncovering a 56-page SOP that turned into a 70-row spreadsheet for one customer alone.
Chapter Markers
00:00 Introduction: capability and trust as one problem
01:07 The agentic AI initiative at Welocalize
03:07 Why rules-based automation hit a technical debt wall
06:45 Decomposing project complexity: common cause vs special cause
10:04 The ethnographic study: watching before automating
12:32 Building the shadow test: 30 days, human vs agent
16:04 Real hiccups: parsing failures and UI friction
19:00 Where human-in-the-loop verification goes next
22:52 Benchmarking against human error, not perfection
26:07 Trust, blame, and why systems feel different to people
29:22 Gathering evidence leadership will actually believe
38:42 Building a repeatable agentic engineering practice
40:50 How AI is changing who can build automation
49:00 Evals-based development as a precondition, not an afterthought
Useful Links & Resources
- Matthew Sekac on LinkedIn: https://www.linkedin.com/in/matthew-sekac-8884894
- Welocalize: https://www.welocalize.com
Connect With the Show
- Host LinkedIn: https://www.linkedin.com/in/b-ross-katz/
- Host on X: https://x.com/brosskatz
- CorrDyn on LinkedIn: https://www.linkedin.com/company/corrdyn/
- Website: https://corrdyn.com
If you're wrestling with the same problem, proving an AI system's reliability before anyone will trust it with production work, we want to hear how you're approaching it. Drop a comment with what your organisation's baseline for "good enough" actually is.
If you want to talk about your data challenges, or you think we got something wrong, find us at corrdyn.com.