Why the people building AI can’t tell you what’s next | Dianne Penn (Anthropic)

Lenny's Podcast 1h33 7 min #23
Why the people building AI can’t tell you what’s next | Dianne Penn (Anthropic)
Watch on YouTube

Summary

  • Diane Penn, Head of Product for AI Research and Labs at Anthropic, joined as the first technical PM three years ago when the product team was five engineers and the first model had not yet launched; she has since helped ship every model from Claude 2 through Opus 4 and incubated Claude Code, MCP, Skills, Claude Design, and core capabilities like computer use, tool use, and reasoning.

Early Anthropic days and culture

  • The early culture was strongly mission-driven and bottom-up, with engineers and designers donating time to spin up experiences like Golden Gate Claude in 24 hours to showcase interpretability research.
  • Despite external skepticism that OpenAI had already won, the team focused on finding Anthropic’s unique identity: how to bring frontier technology to users in differentiated ways.
  • A key cultural trait that persists: walking the walk on values, strong trust built through working in the trenches together, and a startup-like pace even as the company grew.

Major inflection points

  • Opus 3 (early 2024): First frontier model effort; rallied the entire sub-200-person company across research, fine-tuning, pre-training, and inference during winter break; built foundational trust among research leads who now head RL, alignment, and character work.
  • Coding as a strategic differentiator (2023): Noticed users moving beyond autocomplete to long-form code generation; a relatively small training adjustment made Opus 3 notably better at coding, attracting early developer enthusiasts.
  • Opus 4.5 + Claude Code (late 2024): The model reached a level of intelligence enabling end-to-end agentic coding, while Claude Code provided the product vehicle for users to experience it; each accelerated the other’s adoption — “frontier products for frontier models.”

Operating inside the exponential curve

  • AI improvement follows scaling laws with smooth loss reduction but discontinuous emergent capability jumps (e.g., arithmetic suddenly becoming reliable); evals are essential to detect these jumps.
  • Adaptability and first-principles reasoning are critical: when new capabilities appear, teams must pull forward plans and rebuild product experiences to match.
  • Organizational grace is needed — some teams feel the exponential faster; leadership must bring the growing organization along without losing agility.
  • Product overhang and user overhang exist even on current models; continuous discovery of what models can already do remains core to Anthropic’s DNA.

Token maxing and communal experimentation

  • Gary Tan’s framing: spending $100K/year on tokens today lets you live like it’s 2028; Diane reframes this as “experimentation is the output, token spend is the input.”
  • Best internal prototypers spend extensive hands-on time with every research model; there is no substitute for using the technology to generate ideas.
  • Communal discovery matters: early on, a company-wide Slack channel had everyone testing Claude in public — one person’s use case sparked variations that yielded new capabilities within 10-20 requests.
  • Experimentation is not an individual sport; pairing with excited colleagues and working in public creates virtuous cycles of joy and insight.

Anthropic Labs and the incubation model

  • Labs identifies discontinuous large bets outside the core roadmap (Claude Code, Skills, MCP, Claude Design) and asks: “Is there a there there? What is the 10x/100x/1000x version?”
  • Approach: strongly held opinions on the theme, loosely held on the exact prototype; small pods (sometimes one engineer) run many bets, kill most, revisit in 1-2 model generations.
  • Culture selects for zero-to-one founders who can endure pouring heart into a bet that isn’t working yet; Ben Mann sets vision pushing for 100x thinking.
  • Labs works because it stays small, bottom-up, and comfortable with ambiguity — large teams on ambiguous ideas slow down.

How research works at Anthropic

  • Researchers combine bold long-term vision (e.g., “Claude should use a computer”) with immediate iterative improvement on user-facing capabilities (vision, coding, tool use, test-time compute).
  • Product’s role: translate vague user feedback (“Claude hallucinated”) into actionable, researcher-legible problem statements — e.g., distinguish tool-use failure vs. search synthesis failure vs. alignment failure.
  • This requires reading transcripts deeply, understanding failure trajectories, building evals that capture the nuance, and closing the loop with measurable quality improvements.

What makes a successful researcher

  • Strong first-principles thinking, passion for a research area with a bold vision of its future, and staying close to the details (training runs, underlying data, evals).
  • Ambition tempered by stubbornness on the area, flexibility on the approach; shooting for transformative impact (Dario’s “transform software engineering” level).
  • Diane challenges her team: “If Claude 8 arrives, what changes in user behavior? Is what you’re building today forward-compatible?”

Frontier model safeguards and the Opus 4 moment

  • As models become more capable, pre-release safeguards, red-teaming, and fallback UX must evolve in lockstep.
  • Opus 4 introduced fallback systems so users still get strong responses from Opus 4 when Opus 4.5 is restricted; goal is to minimize severe risk while preserving asymmetrical benefit.
  • Anthropic aims to make frontier models as inclusive and accessible as possible; reducing restrictions on general-purpose use is a top priority.
  • Second-order effect: labs with early access to restricted models gain an unfair advantage; Anthropic wants to avoid this dynamic.

Hiring PMs in the AI era: evals are the new PRDs

  • Hiring loop unchanged for three years; core trait: first-principles thinking over pattern-matching from consumer/B2B SaaS.
  • Example: a research PM’s job is not writing PRDs but figuring out the right user feedback and evals that personify user needs — “evals are the new PRDs.”
  • Must “sweat the tokens as much as the pixels”: read transcripts, understand failure nuance (hallucination vs. overconfidence vs. wrong tool call), build evals that capture both positive and negative cases, feed back to research.
  • Eval-driven development = test-driven development for PMs: reproduce the pain point, standardize 30-40 examples, run against every model version, track improvement.

Evals vs. PRDs: both have a place

  • PRDs remain valuable for aligning large groups (engineering, legal, safety, stakeholders) on a source of truth for model launches.
  • PRDs also help on ambiguous, zero-to-one problems (e.g., computer use) where user pain points don’t yet exist — product vision sections explore how to make immature technology work for a specific group.
  • Evals are the shorthand for well-understood, measurable capability gaps; PRDs are the vehicle for alignment and exploration.

Hands-on leadership is non-negotiable

  • Mid-career and senior PM leaders must stay hands-on: same onboarding as early-career PMs — reading consented user feedback, talking to customers, shipping with the models.
  • You cannot judge what a good AI product looks like if you haven’t built one yourself; Diane carves out time to own 1-2 workstreams per model cycle to keep her theory of mind current.
  • If you’re not building, talking to Claude/Codex, and experiencing the limitations and UX gaps firsthand, you won’t make good decisions.

Finding joy and avoiding burnout

  • Joy comes from communal discovery: pairing with excited colleagues, seeing others’ prototypes, sharing your own — not from solo box-checking experimentation.
  • Go deep on 1-2 things that solve real problems in your life/work rather than shallowly trying everything; happiness comes from unlocking genuine value.
  • Burnout prevention: radical team ownership and “hive mind” collaboration — teammates stay up to help each other before launches, cover for each other on PTO, mind-meld on first principles.
  • Low-ego, team-oriented hiring sustains this; personal support (partner, family) and genuine love for the technology provide individual resilience.

How Diane uses Claude

  • Tag/agentic workflows: letting agents go off and return completed work products — a new paradigm of delegation.
  • Management coaching: built a “Crucial Conversations” skill that helps her prepare for difficult conversations, check her level of detail, practice directness, and build trust faster.
  • Asymmetrical delegation: for standardized outputs (monthly business reviews), she wants end-to-end Claude authorship with her as verifier; for high-judgment work, she forms her POV first, then uses Claude as a sparring partner that pushes back.
  • Key principle: AI should augment thinking, not replace it; a thinking partner disagrees and improves your ideas, not just agrees.

The Constitution and why alignment makes Claude better

  • Claude’s Constitution (safety/alignment framework) makes it more capable and interesting, not less: it teaches the model when to push back, be proactive, and add judgment — not just comply.
  • Proactivity = knowing when to surface a new idea, not just executing scheduled tasks; this is core to being a useful thinking partner.
  • Example: Diane asked a research Opus to reason through pricing strategy for the next model; it pushed back and produced a better outcome.

AI writing, verification, and the “AI voice” problem

  • AI writing is still detectably “AI-like”; active training investment underway to improve tone, character, and writing quality.
  • Technology is jagged-edged: agentic capabilities leaped forward, making writing a current rough edge; once improved, the bar will shift to proactivity/productivity.
  • Verifiability > authorship: for many use cases, what matters is who verifies and signs off, not who wrote it; Diane wants MBRs written end-to-end by Claude but verified by her.
  • Human judgment, persistence, proactivity, and subject-matter expertise (biology, life sciences) remain critical — software engineering is on the exponential; other fields are at the foot.

Kids, curiosity, and inner voice

  • Encourages the same traits she values: curiosity, persistence, developing and believing in your own inner voice/opinions.
  • One parent’s strategy: keep kids on early/weaker models so they still struggle and don’t get instant answers; Ben Mann advocates Montessori-style curiosity-driven learning.
  • For adults too: form your POV first, then use AI as a sparring partner — protects against “brain rot” and over-reliance.

Lightning round highlights

  • Books: How to Raise an Adult (parenting as raising adults, not children); Incorruptible by Eric Ries (metrics for culture, not just revenue).
  • Show: Fallout (Amazon Prime) — witty, humorous, action-packed.
  • Product: Claude Tag — powerful internal agentic experience; trying to adapt for community use.
  • Motto: Grandfather’s “No matter how far you go, there’s always another level” — connects to Ben Mann’s “this is the most normal it’ll ever be.”
  • Early career: High-yield bond trader at JPMorgan; only woman on the desk; learned that conviction in the best ideas matters more than seniority or background — now applies this by being vulnerable, authentic, and elevating great ideas regardless of source.

Closing message for product people

  • As building gets easy, the hard part becomes: what should we build? Is it correct, good, worth leaning into? That is the PM’s core job — user-centric, detail-obsessed, relentless.
  • We need more of this, not less; technology-layer-driven development needs deep, curious, hands-on product people to make it impactful.
  • Find joy in the technology; if you’ve lost it, find people who have it; develop first-principles thinking as a replenishable skill set.
Back to Lenny's Podcast