2026-09-28 23:03:13
Listen now on YouTube • Spotify • Apple Podcasts
Brought to you by:
OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
In this solo episode, Claire tests Jev, TypeSafe AI’s new decision model that returns structured choices, scores, and probabilities instead of generated text. She uses it to analyze 1,700 pull requests for 9 cents, map her Claude and Codex usage, triage email, search 4,500 YouTube comments, and process 200,000 classifications for about $4. She also explains why Jev works best alongside a frontier model and how its speed and pricing make entirely new kinds of real-time apps and large-scale analysis practical.
Jev is a decision model, not a language model, and that distinction can make many tasks dramatically cheaper. Instead of generating text, it returns predefined values such as a category, score, or probability. Claire believes this covers roughly 90% of what many software workflows actually need, at 4 cents per million input tokens with no output-token fee.
It cost Claire 9 cents to understand where two years of engineering work went. She used Jev to compare 1,700 ChatPRD pull requests across 17,000 pairs, then had Gemini Flash Lite label the resulting clusters. In about two minutes, she learned that nearly 30% of the company’s engineering work had gone toward platform, security, and infrastructure.
Some of the most useful analysis is already sitting on a local computer. Claude Code and Codex store past sessions locally, allowing Jev to classify them in minutes. Claire discovered that engineering had fallen from nearly 100% of her AI usage in January to less than 40% by September, with agents and media publishing filling the gap.
Jev becomes far more powerful when paired with a frontier model. Claire uses Jev to classify, cluster, filter, and route large datasets, then sends only the most important groups to GPT-6 Astra for deeper reasoning. For ChatPRD’s product insights graph, this approach processed 1,100 signals and completed 200,000 operations for about $4 on the Jev side.
Jev’s pricing changes which ideas are worth building. Because it returns small predefined values instead of generating long responses, TypeSafe charges nothing for output tokens. Claire spent less than $10 on Jev during the week, making classification workloads that would normally be expensive at scale feel almost free.
Jev makes real-time AI loops practical. Claire built a voice app that turns a spoken phrase into a color, matches it with a quote based on sentiment, and displays everything almost instantly. Jev made its decisions so quickly that the quote API became the slowest part of the workflow.
YouTube comment analysis is an immediate use case for any podcast team. Claire classified 4,500 How I AI comments by sentiment, identified 58 containing episode ideas, and built a keyword search that scans the full dataset in under a second. The results showed strong demand for a Grok versus Muse comparison and an 80% positive response to the “Claude Code for product managers” episode.
The real skill is recognizing where a pipeline only needs a decision. Jev will not write documentation or design an interface, but it can sort, route, rank, and filter enormous datasets quickly and cheaply. Claire now asks one question before every build: Where does this workflow simply need to make a decision? That is where Jev belongs.
Jev: AI Data Analysis and Product Insights: https://www.chatprd.ai/how-i-ai/jev-ai-data-analysis-product-insights
↳ Jev GitHub PR Analysis: https://www.chatprd.ai/how-i-ai/workflows/jev-github-pr-analysis
↳ Jev YouTube Comment Analysis: https://www.chatprd.ai/how-i-ai/workflows/jev-youtube-comment-analysis
↳ Jev Multi-Model Product Insights: https://www.chatprd.ai/how-i-ai/workflows/jev-multi-model-product-insights
Listen now on YouTube • Spotify • Apple Podcasts
Claire tests Claude Opus 5.5 after months of leaving Claude out of her daily workflow. She puts it through long-running agentic tasks, frontend prototyping, writing, SVG illustration, computer use, and video editing to see where it earns a place back in her stack. She also shares why she is pairing it with Codex for cross-model code review, where Claude’s safety limits still get in the way, and which tasks remain firmly in Codex territory.
A model’s personality can matter just as much as its intelligence. Claire stopped using Claude for months because its rambling, preachy, and overly verbose replies made it unpleasant to work with. Opus 5.5 is the first model in the family that no longer makes her blood boil, which is a meaningful improvement even if no benchmark captures it.
Opus 5.5’s lower price and faster performance make long-running agent work more practical. It is 40% cheaper than Opus 5, and Claire found it noticeably faster. It successfully completed four complex tasks spanning inbox triage, backend development, research, and computer use, including runs of up to 82 steps from a single prompt.
Silence during long-running tasks creates its own user experience problem. Opus 5.5 sometimes remains quiet for eight or nine minutes, leaving users unsure whether it is still working. It is a reminder that perceived latency matters alongside actual latency, especially when agents run for extended periods.
Opus 5.5 is the strongest frontend designer Claire has tested so far. Its ChatPRD homepage redesign was bold and polished enough that she plans to ship it. The model handles hierarchy, white space, and visual rhythm exceptionally well, though it still struggles with consumer-app aesthetics and defaults to “Claude orange” without direction.
SVG illustration is an unexpected strength of Opus 5.5. It was the only model Claire tested that produced clean, charming, and animatable character SVGs with consistent styling across multiple expressions. The characters remained visually coherent, and their anatomy mostly made sense.
Opus 5.5 has a clear safety posture, and sometimes that means saying no. It refused when Claire asked it to skip testing and push directly to production, and it may route cybersecurity work to Opus 4.8. Whether that feels reassuring or frustrating depends on the workflow, but its boundaries are consistent.
The best use of Opus 5.5 may be as an adversarial reviewer for another model. Claire now has Codex and Opus review each other’s work rather than using one to replace the other. This cross-model loop catches issues either model might miss alone, making the additional cost worthwhile when quality matters.
Computer use and video editing still belong to Codex in Claire’s workflow. Opus 5.5’s ElevenLabs MCP video test produced weak color grading, too few jump cuts, and sloppy overlays. Codex also remains stronger at computer use in her current setup, giving her no reason to shift either category to Claude.
Claude is back, but it has not replaced Codex as Claire’s daily driver. Opus 5.5 has earned a role in pull-request reviews, architecture questions, and frontend development. Codex’s desktop experience, computer use, and workflow integration still keep it in the primary position.
Claude Opus 5.5 Review: https://www.chatprd.ai/how-i-ai/claude-opus-5-5-review
↳ Claude Opus 5.5 SVG Illustrations: https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-svg-illustrations
↳ Claude Opus 5.5 Frontend Prototypes: https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-frontend-prototypes
Listen now on YouTube • Spotify • Apple Podcasts
Claire takes the How I AI bench live to compare GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more across the work she actually does. She blind-scores writing, frontend prototypes, agent personality, and SVG illustrations, with an AI judge helping evaluate backend work, long-running agents, and computer use. She also checks video edits and a 3D Barbie game build. Along the way, she explains why Astra won her heart, Opus 5.5 won her week, and Sol delivered mixed results while remaining a favorite for everyday work.
Expanding the benchmark from two categories to eight changed what Claire could see. The original How I AI Vibe Review focused on PRDs and frontend prototypes. Adding personal productivity tasks like inbox triage, along with backend development, long-running agent tasks, computer use, SVGs, and video editing exposed clear differences between the models Claire preferred for design and those she enjoyed interacting with.
Opus 5.5 returned to Claire’s workflow because of ergonomics, not benchmarks. After repeatedly asking Claude to communicate like a normal person, she found Opus 5.5 concise, clear, and far less irritating. At one point, Claire thought the old frustration had returned, then realized she had accidentally selected Opus 5. The difference was that obvious.
Making Opus 5.5 quieter also made it feel slower, even when it was not. Long stretches of silence can make users wonder whether the model is still working. GPT-6 Sol found a better balance in Claire’s testing, narrating enough to feel responsive without creating additional noise.
GPT-6 Sol’s lower price changes how teams should think about model selection. Learning that Sol costs roughly half as much as Opus 5.5 immediately changed how Claire thought about routing work. She also believes teams should optimize caching before obsessing over model choice, since ChatPRD has seen significant savings when its caches are configured properly.
Dash-heavy writing is an immediate warning sign in Claire’s benchmark. Two models received a 1 out of 5 for agent personality because nearly every message contained an em dash. It may sound overly specific, but Claire sees it as a reliable signal that a customer-facing agent will sound like generic AI writing instead of a natural collaborator.
Claire and the AI judge disagree, which makes the benchmark more useful. The judge favored Fable and rated Sol lower, while Claire preferred Astra. The difference reflects two definitions of quality: the judge rewards correctness and structure, while Claire measures how much she actually wants to use the model.
The blind SVG comparison changed Claire’s earlier verdict. In her standalone Opus 5.5 review, Claire favored its character illustrations. But in this live blind comparison, Astra and Sol came out ahead on character SVGs, surprising her after she had predicted a Claude win.
Opus 5.5 vs. GPT-6 Sol Blind Test: https://www.chatprd.ai/how-i-ai/opus-5-5-vs-gpt-6-sol-blind-test
↳ AI SVG Icon Generation: https://www.chatprd.ai/how-i-ai/workflows/ai-svg-icon-generation
↳ AI Inbox Triage and Email Drafts: https://www.chatprd.ai/how-i-ai/workflows/ai-inbox-triage-email-drafts
↳ Blind Test AI Models: https://www.chatprd.ai/how-i-ai/workflows/blind-test-ai-models
If you’re enjoying these episodes, reply and let me know what you’d love to learn more about: AI workflows, hiring, growth, product strategy—anything.
Catch you next week,
Lenny
P.S. Want every new episode delivered the moment it drops? Hit “Follow” on your favorite podcast app.
2026-09-28 20:03:04
Jev is TypeSafe AI’s new decision model. It returns type-safe structured values (a choice, a score, a probability) instead of generated text, at 4 cents per million input tokens with no output charge. This week I ran it on five real projects: PR categorization, a meta-analysis of my own Claude and Codex sessions, Gmail triage, the ChatPRD product insights graph, and a live audience dashboard built from 4,500 YouTube comments.
Listen or watch on YouTube, Spotify, or Apple Podcasts
What makes Jev fundamentally different from every other model I’ve used
How I analyzed 1,700 PRs for 9 cents and what I found out about where my engineering effort actually went
The personal meta-analysis you can run on your own Claude and Codex sessions right now
Why I stopped using Jev alone, and what I pair it with now
How I turned 4,500 YouTube comments into a searchable audience dashboard for almost nothing
The real-time app I built in an afternoon that shows something surprising about Jev’s speed
Why Jev’s pricing model is different from any LLM I’ve used, and what it makes practical to build
The ChatPRD product insights project: 1,100 signals, 200,000 classifications, and what it cost me
OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
(00:00) Jev launch and what makes it different from every other model
(02:49) Type-safe values explained
(05:28) Understanding Jev outputs
(07:39) Use case 1: PR categorization and pairwise clustering
(11:12) Use case 2: analyzing your own local Claude Code and Codex sessions
(13:00) Use case 3: Gmail triage with Jev scoring and LLM follow-up
(14:30) Use case 4: ChatPRD’s product insights graph
(18:17) Demo: How I AI audience signal dashboard
(22:14) Demo: voice-to-color emotion-mapping app
(25:16) Jev week recap and what’s coming in episode 2
• Jev (TypeSafe AI): https://typesafe.ai
• Vercel: https://vercel.com/ai
• GitHub API: https://docs.github.com/en/rest
• YouTube Data API v3: https://developers.google.com/youtube/v3
• OpenAI Realtime Voice API: https://platform.openai.com/docs/guides/realtime
• Gemini 3.5 Flash-Lite: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite
• API Ninjas Quotes API: https://api-ninjas.com/api/quotes
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].
2026-09-27 20:32:18
Molly Graham is back for round two, and this one is even more powerful. Molly has spent more than 20 years helping organizations and the humans inside them navigate growth and change. She’s held leadership roles at Google, Facebook, Quip, and the Chan Zuckerberg Initiative and is the host of TED’s WorkLife podcast (which she took over from Adam Grant). She also runs Glue Club, a leadership community for senior operators, and writes a popular newsletter called Lessons.
Listen on YouTube, Spotify, and Apple Podcasts
Why Molly’s famous “give away your Legos” career advice no longer holds true in an AI world
The grief, loneliness, and burnout sweeping through the tech industry right now
Why delegating to AI is fundamentally different from delegating to a human
The fear narrative around AI job displacement, and why it’s overblown
Which Legos you should never give to AI
What the best managers are doing right now
WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more
DX—Engineering intelligence for the AI era
• LinkedIn: https://www.linkedin.com/in/mograham
• Substack: https://mollyg.substack.com
• Website: https://glueclub.com
• The high-growth handbook: Molly Graham’s frameworks for leading through chaos, change, and scale: https://www.lennysnewsletter.com/p/the-high-growth-handbook-molly-graham
• Myspace: https://myspace.com
• ‘Give Away Your Legos’ and Other Commandments for Scaling Startups: https://review.firstround.com/give-away-your-legos-and-other-commandments-for-scaling-startups
• Safeway: https://www.safeway.com
• The Odyssey: https://www.imdb.com/title/tt33764258
• What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams): https://www.lennysnewsletter.com/p/building-the-most-ai-pilled-engineering
• Head of Claude Code: What happens after coding is solved | Boris Cherny: https://www.lennysnewsletter.com/p/head-of-claude-code-what-happens
• How to step confidently into the unknown with Manoush Zomorodi: https://mollyg.substack.com/p/new-worklife-episode-how-to-step
• How tech workers are feeling in 2026: a workforce splitting in two: https://www.lennysnewsletter.com/p/how-tech-workers-are-feeling-in-2026
• “I Know It When I See It” Doesn’t Scale with Hilary Gridley: https://mollyg.substack.com/p/worklife-hilary-gridley
• Foo Camp: https://en.wikipedia.org/wiki/Foo_Camp
• Tim O’Reilly on LinkedIn: linkedin.com/in/timo3
• Adam Mosseri: AI is a tailwind for authenticity: https://www.lennysnewsletter.com/p/adam-mosseri-ai-is-a-tailwind-for
• 10 growth tactics that never work | Elena Verna (Amplitude, Miro, Dropbox, SurveyMonkey): https://www.lennysnewsletter.com/p/10-growth-tactics-that-never-work-elena-verna?utm_source=publication-search
• Waymo: https://waymo.com
• AI’s third era: the rise of persistent AI coworkers | Tara Seshan (Product Lead ChatGPT Work): https://www.lennysnewsletter.com/p/ais-third-era-the-rise-of-persistent
• Airbnb: https://www.airbnb.com
• Booking.com: https://www.booking.com
• Brian Chesky’s secret mentor who died 9 times, started the Burning Man board, and built the world’s first midlife wisdom school | Chip Conley (founder of MEA): https://www.lennysnewsletter.com/p/chip-conley
• Hugging Face: https://huggingface.co
• You’re closer to an AI expert than you think with Max Mullen: https://mollyg.substack.com/p/worklife-max-mullen
• OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber: https://www.lennysnewsletter.com/p/openais-head-of-design-this-is-the
• Grok Bot: https://x.ai/news/introducing-grok-bot
• Why Clay Has an AI Writing Policy: https://www.clay.com/blog/ai-writing-policy
• Why Netflix is betting on systems thinkers—not specialists—in the AI era | Elizabeth Stone (CPTO): https://www.lennysnewsletter.com/p/netflix-cpto-on-ai-and-the-future
• The Reverse Centaur's Guide to Life After AI: https://www.amazon.com/Reverse-Centaurs-Guide-Life-After/dp/037462156X
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].
Lenny may be an investor in the companies discussed.
2026-09-27 01:54:02
👋 Hello and welcome to this week’s edition of ✨ Community Wisdom ✨ a subscriber-only email, delivered every Saturday, highlighting the most helpful conversations in our members-only Slack community.
2026-09-23 07:12:49
Listen or watch on YouTube, Spotify, or Apple Podcasts
I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived.
How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it
Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend
Where Sol still wins me over on clear writing, readable PRDs, and price
The character SVG results that completely overturned my prediction about Anthropic
What happened when I asked these models to edit video, and why I think skills explain part of the disappointment
Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t
(00:00) LIVE setup and new model launches
(01:30) What’s new in Opus 5.5, Sol, and Luna
(04:11) Guardrails, personality, and speed
(09:00) The How I AI bench and blind evaluation process
(11:31) Email and personal-productivity results
(13:50) Frontend prototype vibe checks
(24:10) Backend, agent personality, and long-running tasks
(28:25) SVG illustration test
(29:48) AI video-editing results
(30:43) Predictions before the reveal
(31:20) Barbie Bench: the 3D fashion-game test
(34:17) Results: Astra, Sol, and Opus 5.5
(35:04) Writing clarity and creative surprises
(36:51) Why the LLM judge disagreed with me
(37:24) What each model is actually best for
• Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5
• GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/
• Codex (OpenAI): https://openai.com/codex
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].
2026-09-23 03:06:40
Listen or watch on YouTube, Spotify, or Apple Podcasts
I’ve been off Claude for months. Not because it got dumb, but because it got annoying. The rambling, the hedging, the preachy little disclaimers on tasks that didn’t need them. I moved most of my daily work to Codex and I didn’t miss it. Then Anthropic shipped Opus 5.5: 40% cheaper than Opus 5, faster, and with what they’re calling a fundamentally different alignment approach. I ran it for a week across real work, including four long-running agentic tasks, a full ChatPRD homepage redesign, an SVG benchmark, and one very firm refusal, and I’m ready to give you the honest verdict. There’s a lot to like. There are still two things that drive me a little crazy. And there’s one capability I genuinely wasn’t expecting.
Why I walked away from Claude entirely, and what it took for me to come back
The real cost math on Opus 5.5 and why pricing matters more for agentic work than single prompts
What happened when I ran four long-running agentic tasks, including one that tried to manipulate Claude mid-run
Why Opus 5.5 is now my go-to for frontend prototyping, and where it still lets me down
The one capability I genuinely didn’t see coming, and no other model in my stack can match it
The moment Opus 5.5 told me flat-out no, and what that says about where Anthropic’s safety posture actually lands in practice
Where Codex still wins, and how I’m splitting my model stack after a full week of testing
(00:00) Why I stopped using Claude
(01:02) What Anthropic says Opus 5.5 is
(01:54) Cost, speed, and benchmark overview
(03:20) Safety, alignment, and the cybersecurity limits
(05:02) How I AI bench
(05:39) Voice test: is it actually not annoying?
(07:54) Long-running agentic task results
(10:50) Frontend prototyping
(17:23) Writing voice and email
(19:41) SVG illustrations
(20:46) Video editing
(21:42) My verdict: what it’s good at, what it still isn’t
• Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5
• ElevenLabs MCP connector: https://elevenlabs.io/mcp
• Codex (OpenAI): https://openai.com/codex
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email [email protected].