Enterprises have spent the past few years experimenting with generative AI, but the harder question is becoming what happens after a promising pilot. Moving into production means proving that an AI system improves a real business process, delivers measurable value and can keep performing reliably once more users, data and workflows are involved.
That also changes what companies need to measure. Accuracy and model performance remain important, but they sit alongside adoption, productivity, cost, governance and the quality of the decisions or outcomes the system supports.
In this TNGlobal Q&A, Niek van Veen, Vice President for Commercial at Thinking Machines Data Science, answered the questions most relevant to measurement, human oversight, workflow redesign, adoption and the boundaries for agent autonomy. He shares on how enterprises can evaluate AI beyond the proof-of-concept stage, where common measurement approaches fall short, and what organizations should look for as they move AI into day-to-day operations.

In a recent report about Thinking Machines’ partnership with Cebu Pacific, the airline cited major time savings in engineering reference searches, legal contract review and project intake. How were those baselines measured, and how are you checking that the faster outputs remain accurate and useful over time?
We start with the actual job: how long does it take today, what does good look like, and what happens if the answer is wrong? That gives us the baseline and tells us how rigorous the evaluation needs to be.
For something like reference search, we test whether the system consistently retrieves the right information. For work that requires more judgment, you need domain experts evaluating whether the output is actually useful.
The benchmark that matters is whether AI is getting the job done better. A model can score well on a technical benchmark and still fail at the business task you built it for.
And evaluation continues after deployment. The underlying information changes, workflows change, and models change. You need to keep measuring the system against the job it is supposed to do.
In legal contract review, where does AI stop and human legal judgment begin? What kinds of clauses or decisions are explicitly kept outside automated or AI-assisted review?
AI is very good at taking repetitive work off an expert’s plate: finding information, surfacing relevant clauses, organizing a document for review. The lawyer remains responsible for the legal judgment.
That distinction becomes especially important when a decision depends on interpretation, commercial context, negotiation strategy or risk appetite. Those are decisions where you want the expertise of the lawyer.
A good AI workflow should give experts more time for the decisions that actually require their expertise. In legal, the goal is to help lawyers get to the important questions faster, with the relevant information in front of them.
The level of human oversight should ultimately reflect the consequence of getting the decision wrong.
Project intake and triage fell from five days to around one day. How much of that improvement came from AI itself versus redesigning the underlying workflow and decision process?
A lot of the value comes from redesigning the workflow around what AI now makes possible.
We first look at where the five days are going: where information gets stuck, which steps are repetitive, and where people are spending time processing information before they can make a decision. Then we redesign those steps.
You don’t get a five-day process down to one day just by giving people an AI tool. You have to change how the work gets done.
That is why we build these systems closely with the people doing the work. They understand the operational reality; we understand what the technology can do. The biggest productivity gains happen when you bring those two things together.
What did you learn about employee adoption? Were the biggest barriers technical, cultural or related to confidence in the outputs, and what changed after hands-on co-creation with business teams?
Giving people access to ChatGPT is the easy part. Changing how they work is harder.
What we’ve found is that people learn AI best by applying it to a real problem they already understand. So instead of starting with generic prompting exercises, we work with teams on their actual workflows. They see where AI is useful, where it makes mistakes, and where they still need their own judgment.
That builds a much healthier kind of confidence in the technology.
The breakthrough happens when people stop asking, “What can AI do?” and start asking, “What part of my job can I do better with AI?”
Once people can answer that question for themselves, you start seeing adoption turn into new ways of working.
Cebu Pacific says its next roadmap includes agentic applications connected to enterprise data and systems. Which types of actions would you be comfortable allowing an agent to execute autonomously, and where will human approval remain mandatory?
With agents, one of the most important questions is: what authority are you willing to give the system?
Start with bounded tasks where permissions are clear, actions are observable, and mistakes are reversible. You can increase autonomy as you build evidence that the system performs reliably.
The threshold should be much higher when an action affects a customer, safety, sensitive information, financial commitments or an important operational decision. Those are areas where human approval and clear escalation paths become critical.
There is a lot of excitement around what agents can do. For enterprises, the more useful question is what an agent should be allowed to do, under what conditions, and who remains accountable. Getting those boundaries right is what will allow companies to deploy agents with confidence.
Niek van Veen is Vice President for Commercial at Thinking Machines Data Science. He works with organizations across Southeast Asia on data foundations and AI applications, drawing on more than 20 years of digital-transformation experience across Asia Pacific and Europe.
Editor’s note: This Q&A has been lightly edited for clarity and TNGlobal house style. The substance of the interviewee’s responses has been preserved.
Share your perspective: TNGlobal welcomes contributed insights and expert commentary from across Asia’s technology and innovation ecosystem. Submit a contribution for editorial consideration, or explore more conversations in our TNGlobal INSIDER and TNGlobal Q&A and Interviews archive.
Cebu Pacific reports faster legal, engineering workflows after enterprise AI rollout

