Your AI Product Needs Evals
A practical guide to building domain-specific evaluation systems for LLM applications, covering error analysis and iterative testing.
A practical guide to building domain-specific evaluation systems for LLM applications, covering error analysis and iterative testing.
A2UI is a streaming protocol enabling AI agents to dynamically generate and update user interfaces in real time.
Substack profile page for Aakash Gupta, a writer who publishes articles on the platform.
A tweet by David Beyer (@dbeyer123) on X, though its specific content couldn't be retrieved from the page metadata.
A tweet by Nate.Google sharing a claimed method for getting brands to rank quickly in ChatGPT and other LLM responses.
A Substack profile page for Rich Holmes featuring articles and newsletters published on the Substack platform.
Newsletter piece exploring how AI tools like Claude Code and Figma are reshaping product design workflows and the growing importance of product taste.
Podcast interview with designer Xinran Ma detailing the workflow PMs can use to go from idea to AI-generated prototype quickly.
Article providing definitions and explanations of AI-related terminology and concepts.
A GitHub repo providing techniques and reference numbers for estimating system performance from first principles.
Personal site of jackdoe, a software engineer, featuring technical notes and projects related to search and information retrieval.
Video interview with AI researcher Yann LeCun discussing his perspectives and reflections on artificial intelligence's development and future.
A step-by-step playbook for tackling analytical thinking questions in product manager interviews, covering frameworks and practice strategies.
Podcast interview with Nesrine Changuel (Spotify, Google, Skype) outlining a 4-step framework for building delightful products.
Article exploring the concept of agency and its applications.
A guide comparing prompt engineering, RAG, and fine-tuning to help product teams choose the right approach for building custom AI features.
A free hands-on tutorial teaching product managers to use Claude Code's AI agents and file operations for PM workflows.
Explains the demand-driving-supply growth loop used by marketplace giants like Uber and Airbnb to bootstrap and scale supply.
Curated article collection of important readings for product managers and builders.
Article recommending Claude Code tool for product managers and developers.
LinkedIn article explaining first principles thinking as a problem-solving method, using Elon Musk's innovation approach as an example.
Blog post by Paul Buchheit on product quality and user experience priorities.
Video of Ira Glass discussing the importance of finishing creative work despite early flaws, a key storytelling lesson for product builders.
Archive collection of Marc Andreessen's product management essays and insights.
An article outlining the Four Pillars of Integrity framework from the Conscious Leadership Group's resource library on personal and professional growth.
Website of Silicon Valley Product Group offering product management resources.
Article on strategies and approaches for effective learning.
A Reforge article outlining four growth frameworks—covering acquisition, engagement, and monetization—used to scale products to $100M in revenue.
Craft article describing how analytics integration into development environments improves workflows.
Video interview with prompt engineering researcher Sander Schulhoff on which 2025 prompting techniques actually improve LLM outputs.
Podcast interview with Learn Prompting/HackAPrompt founder Sander Schulhoff on current effective prompt engineering techniques versus common misconceptions.
Article on using data-driven methods and structured prompts for Midjourney image generation.
OpenAI's official guide for prompting GPT-5 models.
AI Engineer newsletter and podcast featuring interviews with leaders like Karpathy and Hotz on building agents, models, and infra.
Codecademy course teaching deep learning model development using PyTorch.
Weaviate blog article explaining chunking strategies for RAG pipelines and their impact on retrieval quality and agent memory performance.
Explains why LLM inference is nondeterministic even at zero temperature and how batching/kernel effects cause this, with fixes for reproducibility.
Video explaining fundamental concepts and mechanics of how AI systems work.
Article with instructions and guidance on building a RAG-based chatbot.
1990 NeurIPS academic paper on neural networks and their applications.
Deep dive with Cursor cofounder Sualeh Asif on scaling infrastructure to 1M+ QPS and billions of daily AI code completions.
A community video library where users showcase computer vision projects built with LandingLens, offering practical AI engineering examples.
Guide for training small language models efficiently on Hugging Face.
Stanford's CS336 lecture series teaching how to build large language models from scratch, covering data, architecture, training, and systems.
GitHub repository collecting system prompts and model information from various AI tools.
Video overview of UFO, a UI-focused AI agent that autonomously controls Windows applications by interpreting screen elements to complete tasks.
Landing AI blog post introducing VisionAgent, an agentic system that chains code generation and tool use to solve complex visual reasoning tasks beyond standard VLM capabilities.
A paper introducing PSI, a video world model trained on 1.4T tokens that iteratively extracts and integrates structures like depth and flow for improved prediction and control.
GroqCloud playground lets developers test and run large language models via Groq's high-speed inference API.
Platform for tracking, analyzing, and evaluating LLM prompts and outputs.