How The Product Role Is Moving To Building And Verification
As AI speeds up execution, product leaders must focus more on product verification, evidence, judgment, and what is worth shipping.
TL;DR: AI has become so much more capable that it can handle long-horizon planning and execution. Not only that, it has made product execution much faster than ever before. Earlier, product leaders would spend time planning and then coordinating what needs to be done. They would arrange resources and engineers and keep an eye on the economics of the entire product operation. After AI, execution becomes readily available to everyone; the main concern is whether that execution is worth shipping. This blog explores how the product leader’s role has shifted from execution to product verification and what it means to create artifacts to generate evidence for product growth.
Building Is Getting Faster Than Product Decisions
We already know how organizations, companies, startups, and AI labs are continuously deploying new features on a weekly basis and sometimes on a biweekly basis. As a product leader/builder, the idea is to be as fast as possible and compete with other brands so your products can stand out and reach the audience first. One thing happening in this AI era is that we are building products faster, but we are not sure whether the decisions we made toward that particular feature or launch are good enough to meet user expectations. If user expectations are met, then the revenue and product will grow eventually.
With AI, the product managers can do a wide range of things, like,
Creating prototypes. They can take their imagination or a simple idea and create a prototype out of it.
They can analyze the market. They can analyze what a group of people is missing in their product and target them. They can analyze the feedback. They can analyze how users interact with the product and so much more.
They can take the feedback from the users and create a better feature.
They can create a functional workflow of a product.
They can create an enhanced product architecture, define how the next iteration should look, how users should navigate, remove any constraints that might be blocking users from being more productive, and so much more.
They can essentially code by themselves. If they find something really interesting, or an idea is very practical, they can vibe code and build the entire prototype by themselves to test.
As such, Anthropic analysis shows that there are roughly 400,000 Claude Code sessions each day. And the interesting thing is that users made about 70% of the planning decisions, while Claude Code made about 80% of the execution decisions.
What does that mean?
Essentially, it means that execution remained one of the bottlenecks in the pre-AI era. Now, with AI, all you have to do is spend an ample amount of time planning. The execution part, or more specifically the implementation part, can be handled by AI agents with long-horizon execution. This chart also shows the clear distinction between what to do and how to do it.
But the thing to understand is that AI allows product leaders to build products and their features faster than their capacity to make good product decisions.
So, what is lacking here that could solve that issue?
It is the builder-verifier role.
Let’s explore further.
Product Management Was Built Around Scarce Execution Era
Before the AI era, building a product generally required thorough market research, including finding gaps in the market and understanding users’ requirements, pain points, and solutions. It also included where the market is currently, meaning whether the market will be friendly enough to welcome this new product and who the possible ICPs are.
With all that research, it was important to define:
How the product will look?
What features it should have to cater to the specific ICPs?
The design of the product.
The UI/UX and how it should interact with the user for a better experience.
How should the onboarding look?
What other color designs and schemes would attract users and make them come back again and again?
The other thing was engineering. This means the architecture of the entire product: what stack to use, what language to build in, and everything required to build a good database with a server to use, and things like that. Then there is quality assurance, and finally, the product is released or generally available.
The PM’s role in all these things is coordination. Why coordination? Because execution was not readily available, meaning everything was done by hand manually. There was less automation for coding, so the PM has to ensure:
Which problem needs more time?
Which problem needs to be addressed first and executed?
Which problem deserves more investment, in terms of time and resources, as well as economic investment?
Translate the customer’s problem into a plausible solution.
Gather the team: Engineers specialized in databases, engineers specialized in frontend engineering, engineers to code the entire architecture, etc.
Manage the trade-off: What should be the best possible release for this season to solve the problem for users immediately, and what release could be added later on?
These were some of the things that the PM would do. It is also important to say that the product manager or the product builder’s role was built on scarce execution.
So what changed now? The obvious answer is execution.
AI has clearly paved the way for faster execution, and it has helped PMs, engineers, and designers step into other roles or switch roles. Let’s look at some data from OpenAI analysis. It found that 43.5% of the messages received involved tasks associated with other domains.
Meaning: Product Leaders are looking for ways to code, Engineers are looking for ways to design, and Designers are looking for ways to research and communicate better ideas and decisions.
Previously, the product leaders’ workflow was to generate or work on an idea, then get the requirements, hand off the requirements to specific teams (designers, engineers, etc.), review, and then release. Now, with AI, the workflow has transformed. It is:
Generate an idea, work on it yourself, or use AI to brainstorm and refine it.
Create artifacts and prototypes.
Generate evidence of how it is performing within a small set of users.
Make decisions; fine-tune it before the final release.
The New Bottleneck
With every new advancement and open door, there are always obstructions and hindrances that stop you from moving forward. With AI, it is generating more outputs than what we could handle. AI is helping us to generate more code and artifacts, and with that we can essentially have multiple experiments running in parallel at a given time.
We are creating metadata and artifacts in a huge volume, and the idea now remains:
How can we evaluate all the artifacts and the metadata? Which one is good?
Which one should be exposed to the public?
Which one should be released to the public?
Which one shouldn’t be?
What are the criteria for the release?
What are the criteria to stop or postpone the release and reiterate the product?
It all boils down to evaluation and better decision-making.
A study from Microsoft found that people who have incorporated AI into their workflow had merged roughly 24% more pull requests. This is a great number. But they also warned that the merge was not properly aligned with the output and product values. They were proxies for the output, which essentially means it’s not equivalent to a fully functional product.
It is quite evident that the new bottleneck is idea to output, but output with confidence.
Product Manager Builds Evidence
Product leaders and product managers have knowledge about all the components and nitty-gritty details that a product encapsulates. It can be:
APIs
AI models and their behavior. How to build an agentic stack for the product?
Data and tool cost.
Permissions and guardrails.
Rubrics for evals, etc.
But the thing is, they are not the sole owner of the product. Because they have a wide spectrum of knowledge available to them, they can direct the product and align it with the user preferences. They can also recommend to teams what to do and how to do it, what the approach should be, and what a better UI design or architecture would be for a certain feature.
As such, product leaders can create small, useful artifacts or prototypes that can challenge an assumption or inform a decision. They can then use this information to communicate with designers, engineers, and other stakeholders within the leadership team.
This blog from Anthropic mentions that users using Claude Code made task-specific demonstrations to convey their concerns, issues, appropriate verification requests, growth projections, etc. This proves that a product manager’s job is becoming more about finding truth and underlying patterns that can lead to failure or success.
The essence of AI is that it enhances domain judgment and doesn’t diminish human intelligence and experience.
Anthropic analysis also shows success rate in coding for people in different domains. Using Claude Code, many have obtained the same level of success as a software engineer would. This implies that users can create evidence-based prototypes and presentations through data and code to show whether a product will pass or fail, or the challenges the product can have in the next release.
Verification is Essential to Build a Meaningful Product
Verification does not mean checking the model for the correct output or a plausible answer. It is evaluating whether the entire system meets product expectations. So what does “a system” in this sentence mean?
Essentially, “system” is the complete AI product or workflow being verified. Not just the model itself, not just the output, and not even the permission, guardrails, tools, or API alone. It is the entire system. Because all the components contribute to the output.
Here are five ways in which the entire system can be verified:
Problem verification: Here you should ask whether the problem or the issue is frequent and important enough to solve instantly, or is it painful, meaning is it the cause of the user declining?
Behavioral verification: Is the system able to capture patterns and handle edge cases while completing the task, aligning with the goal, and not steering away or hallucinating from the goal?
Value verification: This is quite simple. Is the product improving customer experience or business outcome?
Operational verification: Is the product or the system reliable, trustworthy, economically and cognitively affordable?
Strategic verification: Is the product worth owning and supporting?
One good practice I learned from OpenAI and Anthropic blogs is to always write evals based on the PRDs. Your evals should not be something that comes out of your imagination. But they should explicitly state product expectations. The eval should know the system purpose, the desired outcome, the major decision points, etc. Also, having a golden set that represents expert judgment is something that I would highly recommend.
Also, a good approach is to write evals based on the 20 or 50 real examples of real failures, and then check them manually without constructing a large benchmark. Lastly, you should incorporate user issues and feedback into the evaluation.
What Product Teams Need and Closing
In this blog, we learned the importance of product verification and how product managers’ roles have transitioned from just coordinating the entire product roadmap to essentially being a part of different aspects of the product development. We gave evidence of how the product might succeed or fail and what issues a product can face. Product verification is not about how well the model works, but it is about how the entire system works. This includes:
The API call
Tool calling
The configuration of the model and model stack
Evals and rubrics
Guardrails, etc.
Before we close, let me point out some of the important things that you can keep in mind when you are verifying the product:
Extend the PRD; it is important that the product PRD is hyper-specific. This is also important because AI models need a well-defined product intent so that the product can be tested frequently to ensure that the product doesn’t drift away.
Also spend time exploring and building new artifacts, such as:
Writing eval cases where the product might fail, which can be done by using AI, to explore where the product might succeed and where the product might fail. It can also highlight issues that are still hidden in production.
What is the acceptance rubric?
What are the guardrails that need to be written?
What are the failure taxonomies?
How long is the model taking to give an output?
What is the user requirement?
Understanding and studying the production traces explores new patterns and limitations of the product.
Introduce a new matrix as a product builder. This is your job: to understand the product and come up with a new matrix to evaluate and verify your product.
That is to say that product verification is an important skill that product leaders should have in this AI era.





