PM in AI-Native Companies: New Skills, New Risks
Lesson 84: PM in AI-Native Companies: New Skills, New Risks
Lesson 84: PM in AI-Native Companies: New Skills, New Risks
Module 7 gave you the Ownership Zones Model (Lesson 65) for assigning responsibility around a model's output, and the Discovery Frontier (Lesson 66) for reasoning about a recommender's explore/exploit balance. Lesson 81 established that some automated decisions legally require a human in the loop, and Lesson 82 established that data handling must be minimized and purpose-limited. An AI-native company — one whose core product is built around large generative models rather than using a model as one feature among many — inherits every one of these concerns simultaneously, and adds a genuinely new one: the model's capability and reliability are not fixed properties you can fully test once and trust forever. They shift with every model update, every prompt change, and every new edge case a large, open-ended user base discovers.
A PM new to AI-native product work tends to treat a generative model the way they'd treat any other software dependency: build it, test it, ship it, move on. This instinct fails because a generative model doesn't produce a fixed, enumerable set of outputs the way traditional software does — it produces a probability distribution over an effectively unbounded output space, meaning "testing it" can never mean the same thing "testing it" means for deterministic software. The central discipline this lesson introduces is refusing to treat model capability as a single number ("is it good enough to ship") and instead asking a two-part question: how capable is the model at this specific task, and how reliably does it perform at that capability level, since these two properties can diverge sharply and require entirely different product responses.
Learning Objectives
- 1
Explain why a generative model's quality cannot be reduced to a single "is it good enough" judgment.
- 2
Apply the Capability-Reliability Matrix to determine the appropriate product response for a given AI-powered feature.
- 3
Identify why evaluation ("evals") must be an ongoing practice rather than a one-time pre-launch gate for AI-native products.
- 4
Explain why per-inference cost economics differ from traditional software's near-zero marginal cost assumption.
- 5
Evaluate a proposed AI-powered feature for whether its automation level matches its actual position on the Capability-Reliability Matrix.
This lesson assumes the Ownership Zones Model from Lesson 65, the Discovery Frontier from Lesson 66, and the Regulatory Surface Map's human-in-the-loop concept from Lesson 81, since AI-native product decisions draw on all three simultaneously.
Capability Is Not a Single Number
Capability Is Not a Single Number
A generative model can be highly capable at a task — able, in its best outputs, to perform at a genuinely impressive level — while being unreliable at that same task, meaning its performance varies significantly across similar-seeming inputs, sometimes producing an excellent result and sometimes producing a confidently-stated but wrong one (a "hallucination"). Capability and reliability are distinct properties, and a product decision informed only by capability ("look how good this can be") without accounting for reliability risks shipping a feature that performs impressively in a demo and unpredictably in production.
The Capability-Reliability Matrix
The Capability-Reliability Matrix
This lesson introduces the Capability-Reliability Matrix:
A task landing in quadrant A can reasonably be fully automated. A task in quadrant B — the most common position for many current generative use cases — calls for keeping a human explicitly in the loop, per the Ownership Zones Model's Zone 4 discipline and, in regulated contexts, per Lesson 81's legal human-in-the-loop requirements. Quadrant D should not ship at all for that specific use case, regardless of how compelling a demo made it look.
Evals as an Ongoing Practice
Evals as an Ongoing Practice
Evals — structured evaluation suites measuring a model's performance on representative tasks — must be run continuously, not once before launch, because model behavior shifts with every underlying model version update and every prompt or system change, directly echoing the Metric Provenance Chain's insistence (Lesson 64) that a metric must be continuously validated, not validated once and trusted forever.
Per-Inference Cost Economics
Per-Inference Cost Economics
Unlike most software, where serving one more user costs almost nothing, generative AI features carry a genuine, often significant per-inference cost that scales directly with usage — a cost structure a PM must model explicitly, since a feature can be capable, reliable, and popular, and still be a poor product decision if its unit economics don't work at scale.
Why AI Products Break Traditional QA Assumptions
Why AI Products Break Traditional QA Assumptions
Traditional software QA rests on a quiet assumption that most PMs never have to state explicitly: a given input, run through a given version of the code, produces the same output every time. This determinism is what makes a fixed test suite meaningful — if a test passes today, it will pass tomorrow unless the code changes. Generative models violate this assumption at the root. The same prompt, run twice against the same model version, can produce two different outputs, and a test suite that passed cleanly last week can begin failing this week with no code change at all, purely because the underlying model provider shipped a silent update, or because the specific inputs users send have drifted from the inputs the team originally tested against. This means a PM cannot treat "we have a test suite and it's green" as evidence of ongoing quality the way they could for deterministic software — the test suite itself must be re-run continuously against production-representative inputs, and its results must be watched for drift, not just checked once at a release gate. This is the deeper reason evals function differently from unit tests, and why an AI-native PM's mental model of "quality assurance" has to be rebuilt rather than simply carried over from prior software experience.
Common Mistakes to Avoid
Treating "the model is impressive in a demo" as sufficient evidence to ship at full automation
A demo showcases a model's best-case behavior, often on inputs the team implicitly selected because the model handles them well. Demos showcase capability, not reliability across the full range of real inputs a production system will actually encounter, including edge cases and phrasing the team never thought to try. A PM who greenlights full automation based on a demo alone is confusing "this can be impressive" with "this will be dependably correct" — the two separate axes the Capability-Reliability Matrix is built to keep apart.
Running evals once before launch and treating the model as permanently validated
Model updates and prompt changes require continuous re-evaluation, because a green eval result today says nothing about tomorrow's underlying model version, upstream provider update, or drift in the inputs users actually send. Treating a passed eval suite as a permanent credential, rather than a snapshot that expires the moment anything upstream changes, is precisely the deterministic-software assumption this lesson argues generative AI breaks. An eval suite for an AI feature has to be re-run against production-representative inputs on an ongoing basis, not checked once at a release gate and set aside.
Ignoring per-inference cost until scale reveals it as a problem
Unlike most traditional software, where serving one additional user costs almost nothing, a generative AI feature carries a real per-inference cost that scales directly with usage. Unit economics should be modeled before launch, not discovered after, because a feature that is capable, reliable, and genuinely popular can still turn out to be a poor product decision if its cost per use erodes or exceeds the value it creates at scale. Waiting until a feature has already succeeded to check whether it can afford to keep succeeding is a needlessly expensive way to learn this.
Confusing a model's stated confidence with actual reliability
A hallucinated answer is frequently delivered with the same fluent, confident tone as a correct one, so a model's apparent certainty is not evidence of its actual accuracy. Teams that read fluency as a reliability signal, rather than measuring actual correctness against representative test cases, can badly overestimate how trustworthy a feature is in production. This is one reason the Capability-Reliability Matrix insists on separately verified reliability data rather than an impression formed from how the output reads.
Treating a feature's Capability-Reliability Matrix quadrant as permanent
The Capability-Reliability Matrix
Apply it by asking: (1) How capable is the model at this specific task? (2) How reliably does it perform at that level across representative real inputs, not just demo inputs? (3) Does the current automation level match the quadrant, or does launch pressure push toward more automation than reliability supports?
Key Takeaway: How will you apply "The Capability-Reliability Matrix" when evaluating trade-offs in your product decisions?
Ready to test your product judgment?
Take the interactive practice quiz for Lesson 84 and build your skill radar dashboard.