What happens when an AI agent's mistake can cost you the client, not just the demo.

You've proven your AI agent works in a sandbox. Getting the business to trust it in production, with real customers watching, is a different problem entirely.

Matthew Sekac is Head of Data & Analytics at Welocalize, where he has spent the past 18 months building an AI agent framework to handle the project management complexity behind thousands of translation and localisation projects. He joined Welocalize in 2020 and has over a decade of senior sales strategy experience in IP and life sciences translation before moving into data leadership.

Matt walks through exactly how his team proved an AI agent framework could handle over a million manual tasks a year without a mountain of rules-based technical debt. You'll get a concrete method for testing agent reliability against human benchmarks before you ever touch production and a way of thinking about human-in-the-loop verification that scales instead of adding new work.

This one covers ethnographic research into how project managers actually work, shadow testing agents against human output, and the evals-based thinking needed to get leadership buy-in for AI at scale. It's built for data and analytics leaders trying to move AI agents from pilot to production, not for anyone looking for a quick AI hype fix.

Key Takeaways

- Welocalize ran a 30-day shadow test where an AI agent mirrored every human decision on live tasks before a single customer saw it, and that proof, not confidence, is what got leadership to say yes.

- The team found their "quality check" agent was inversely correlated with actual accuracy: the more confident it was that a human wouldn't override it, the more likely a human did. The reason why says a lot about where reasoning agents still fail.

- Nobody asks if a human process is error-free before automating it, but everyone expects zero errors from AI. Matt argues that a mismatch is the real reason AI trust conversations go sideways.

- Before writing a single rule, the team spent weeks just watching project managers work, uncovering a 56-page SOP that turned into a 70-row spreadsheet for one customer alone.

Chapter Markers

00:00 Introduction: capability and trust as one problem

01:07 The agentic AI initiative at Welocalize

03:07 Why rules-based automation hit a technical debt wall

06:45 Decomposing project complexity: common cause vs special cause

10:04 The ethnographic study: watching before automating

12:32 Building the shadow test: 30 days, human vs agent

16:04 Real hiccups: parsing failures and UI friction

19:00 Where human-in-the-loop verification goes next

22:52 Benchmarking against human error, not perfection

26:07 Trust, blame, and why systems feel different to people

29:22 Gathering evidence leadership will actually believe

38:42 Building a repeatable agentic engineering practice

40:50 How AI is changing who can build automation

49:00 Evals-based development as a precondition, not an afterthought

Useful Links & Resources

- Matthew Sekac on LinkedIn: https://www.linkedin.com/in/matthew-sekac-8884894

- Welocalize: https://www.welocalize.com

Connect With the Show

- Host LinkedIn: https://www.linkedin.com/in/b-ross-katz/

- Host on X: https://x.com/brosskatz

- CorrDyn on LinkedIn: https://www.linkedin.com/company/corrdyn/

- Website: https://corrdyn.com

If you're wrestling with the same problem, proving an AI system's reliability before anyone will trust it with production work, we want to hear how you're approaching it. Drop a comment with what your organisation's baseline for "good enough" actually is.

If you want to talk about your data challenges, or you think we got something wrong, find us at corrdyn.com.

Podden och tillhörande omslagsbild på den här sidan tillhör CorrDyn. Innehållet i podden är skapat av CorrDyn och inte av, eller tillsammans med, Poddtoppen.