Why AI is going vertical (again) | Dianne Penn (Anthropic)
Key insights
Books referenced
- How to Raise an Adult - Julie Lythcott-Haims - Penn's go-to personal recommendation on what it actually means to raise capable adults rather than sheltered children.
- Incorruptible - Eric Ries - Recent Audible listen; Penn liked its framing of how great companies sustain their values through the metrics they choose, not just revenue.
- Crucial Conversations - Kerry Patterson, Joseph Grenny, Ron McMillan, Al Switzler - Penn built a Claude skill around this book's framework to help her prep for difficult management conversations in the moment.
Media referenced
- Eric Ries episode of Lenny's Podcast - podcast - Referenced as the companion listen to Penn's take on 'Incorruptible' and building enduring companies.
- Fiona Fung episode of Lenny's Podcast - podcast - Fung (manager of Claude Code and Cowork) suggested the burnout question Lenny asks Penn, and is cited on software engineering feeling lonelier now that teams work with agent fleets.
- Ben Mann's episode of Lenny's Podcast - podcast - Referenced twice: Mann's line that 'this is the most normal it's ever going to be' and his answer about curiosity and Montessori schooling for kids.
- Andrew Ambrosino episode of Lenny's Podcast - podcast - OpenAI's Codex app lead, cited as independently agreeing with Penn that PRDs are not dead.
- Fallout - show - Penn's favorite recent watch, an Amazon Prime series based on the video game, binged during time off.
Companies
- Anthropic - Penn's employer since 2023; the episode's central subject, covering its product and research culture.
- OpenAI - Referenced repeatedly as Anthropic's early competitive benchmark and via its Codex product lead.
- WorkOS - Episode sponsor; enterprise-readiness APIs (SSO, SCIM, RBAC, audit logs).
- Mercury - Episode sponsor; banking platform Penn's host uses, highlighted for its new 'Command' AI interface.
- JPMorgan Chase - Penn's employer before tech, as a high-yield bond trader on the trading floor.
- Y Combinator - Gary Tan (YC president) is credited with the 'token maxing' framing discussed at length.
- Amazon - Penn's employer immediately prior to Anthropic.
Techniques and frameworks
- Evals are the new PRDs - Anthropic's research product management team defines and proves user value through eval sets built from real failure transcripts, not written product requirement docs.
- Token maxing - Gary Tan's framing, discussed and reframed by Penn: spending heavily on tokens today to live like a 2028 user, though Penn recasts the real goal as experimentation volume, not spend itself.
- Golden Gate Claude - A 24-hour 2024 interpretability demo where a dialed-up 'Golden Gate Bridge' feature made Claude reference the bridge in every response; an early example of turning research into a shipped user experience fast.
- Labs incubation model - Anthropic Labs holds a strong opinion about the problem area and a loose one about the exact prototype, runs small pods (sometimes one engineer), and treats a failed bet as valid learning to revisit a model generation later.
Summary
Dianne Penn, Anthropic's first technical product manager and now head of product for its AI research and labs teams, walks through the company's product history from a five-engineer team in 2023 to shipping "more than a year's worth of 2024 model releases within a single quarter" today. She frames two inflection points as the real turning points: training Opus 3 to write long-form code rather than just autocomplete it, and Opus 4.5 arriving alongside Claude Code, since neither the model nor the product would have had its adoption moment alone. She also recounts Golden Gate Claude, a 24-hour 2024 stunt where an amplified interpretability feature made Claude reference the Golden Gate Bridge in every response, as the early moment the company found its identity for turning research into fast, distinctive shipped experiences.
The core of the conversation is how product management itself has changed at Anthropic. Penn's central claim is that "evals are the new PRDs": instead of writing product requirement docs from user interviews, her team reads failed model transcripts to pinpoint the actual failure mode (a missed tool call, a knowledge retrieval miss, an alignment issue) and encodes that as a measurable eval set researchers can act on. She likens this to test-driven development for PMs, and connects it to how emergent capabilities work: scaling laws produce smooth loss curves, but specific abilities jump into existence suddenly and unpredictably, which is why evals rather than intuition are the tool for catching both product opportunities and safety risks. She's explicit that PRDs aren't dead, especially for aligning large stakeholder groups on ambiguous, not-yet-validated ideas like early computer use, but for well-defined problems the eval has become the shorthand artifact.
Penn describes Anthropic Labs' incubation model as holding a strong opinion about the problem area and a loose one about the exact prototype, run by small pods (sometimes a single engineer), where a failed bet is treated as valid learning to revisit a model generation later rather than a wasted effort. On hiring, she says her team's PM evaluation criteria haven't changed in three years and center on first principles thinking over pattern-matched experience from prior consumer or SaaS roles; she applies the same onboarding plan to senior tenured hires as early-career ones, and expects managers to personally ship with the models, not just direct people who do, because you can't recognize a great AI product without having built one yourself.
A recurring theme is Anthropic's explicit design choice to make Claude push back rather than simply agree, which Penn frames as central to both alignment and usefulness: a model that just agrees makes a user's thinking worse, while one that challenges a half-formed idea makes them better. She personally splits her own Claude usage by how much judgment she wants to preserve, forming her own point of view before bringing Claude into higher-stakes writing but delegating routine work like monthly business reviews end to end and acting as reviewer rather than author. She also names writing as one of the current jagged edges of frontier models, attributing it to training sequencing (agentic tool use came first) rather than a knowledge gap, and says it's now an active investment area.
The episode closes with a lightning round covering her book recommendations (a parenting book on raising capable adults, Eric Ries's "Incorruptible" on sustaining company culture through the metrics you choose, and Crucial Conversations, which she turned into a personal Claude coaching skill), her view that durable human value sits in judgment and persistence as capability keeps expanding, and how she's raising her own kids toward curiosity and trusting their own inner voice rather than leaning on AI tools early. She credits Anthropic's low-ego, team-oriented hiring and a strong sense of shared ownership (rather than any individual heroics) as what has kept her from burning out through a year where the company shipped more models in one quarter than it did across all of 2024.
Notable Quotes
"We actually have a saying on the team of evals are the new PRDs." - Dianne Penn
"You have to sweat the tokens as much as you sweat the pixels." - Dianne Penn
"You should come away at the end of the day having better ideas because you've worked with Claude. That should be the hero goal, not just making your ideas 10% better." - Dianne Penn
"No matter how far you go, there's always another level." - Dianne Penn, quoting her grandfather's life motto
"It's not just about raising the IQ of experiences we build - it's about having better conversations with each other, being better managers." - Dianne Penn