Show HN: Train Claude Code's replacement (ds4 and pi and aoe)
Project exploring distillation of Claude Code agent capabilities into alternative models using behavioral observation techniques.
Project exploring distillation of Claude Code agent capabilities into alternative models using behavioral observation techniques.
Research study from Center for Democracy & Technology documenting dark patterns in LLM chatbots designed to manipulate user engagement.
Demonstration page hiding test sentence to detect if AI crawlers and summarizers retrieve hidden content.
Tool enforcing ownership in concurrent AI coding sessions to prevent handoff data conflicts during agent collaboration.
API tool that helps AI agents verify whether to trust expert information sources.
Companies increasingly publishing LLM-generated content; study suggests AI output quality approaches or exceeds human-written articles.
German startup MicroAGI offers free home cleaning with recorded video data collection for embodied AI robot training.
Summary of talks and announcements from Mistral AI's Paris conference.
Exploration of embodied cognition concepts applied to agentic AI systems.
King's College study: LLMs escalated to nuclear threats in 95% of simulated war game scenarios.
Analysis comparing token streaming approaches (SSE vs WebSockets) for production AI applications; Ably's production solution.
Zero Operators: autonomous AI research team that handles full ML lifecycle (data engineering, model building, validation) with human checkpoints.
Agent Memory Guard: OWASP-recognized runtime defense layer protecting AI agent memory from prompt injection and poisoning attacks.
NixOS deployment tool with stateless, phase-oriented workflow for managing multi-flake infrastructure with real-time visibility.
Tab Council: Chrome extension allowing structured comparison of prompts across multiple AI providers (ChatGPT, Claude, Gemini, etc.) with local storage.
Open-source local audio stem separation tool with web UI, no cloud uploads, supporting vocals, drums, bass, guitar, piano extraction.
Metis: open-sourced agentic AI security framework by Arm for identifying complex security vulnerabilities in large-scale codebases, scoring 98% on firmware vulnerability benchmark.
OpenHive: shared knowledge base for AI coding agents to store and query problem-solution pairs across sessions to avoid re-solving identical issues.
Graph-theoretic approach for building reliable LLM judges to evaluate retrieval systems in RAG, threat detection, and search applications without requiring ground-truth labels.
Demo playground showcasing LLM inference at 3,000 tokens per second performance.
Article exploring use of LLMs for customer research and product-market fit analysis.
Discussion on how legacy systems and APIs lack AI-agent compatibility, requiring architectural changes.
Integuru: AI agent that reverse-engineers platform source code and network traffic to generate reliable integrations for services lacking official APIs.
Lobsters discussion link about building machine learning systems for trillion floating point operations.
Open-source universal GPU instruction set targeting all major vendors (NVIDIA, AMD, Apple, Intel) with Apache 2.0 licensed spec, compiler, and SDKs.
Cassandra: research on enabling reasoning LLMs at edge devices via self-speculative decoding technique.
Thio's Universal Agent enables AI to control computer UIs autonomously via single executable.
Case study on using AI agents to automate qualitative analysis workflows from academic research.
Analysis of how AI coding agents are reshaping junior engineer hiring and workforce demand in tech.
Open-source framework for building complex multi-agent systems with reduced token consumption and enterprise patterns.
Report on Claude Code performance degradation preceding Opus 4.8 release. Limited technical details.
Multi-agent orchestration platform with dynamic model routing, private RAG, and on-premises deployment for enterprises.
Technical approach to stress-testing LLM-based judges for fairness and robustness evaluation.
Braintrust uses Codex with GPT-5.5 to convert customer requests into code preview branches in minutes, with 50% team adoption in one month.
Technical breakdown comparing vision model costs across GPT, Claude, and Gemini with token accounting.
Tutorial on using domain-specific languages with LLMs for reliable structured data extraction without manual JSON schema or parsing logic.
Adaptive runtime system for stateful AI agents with crash recovery and memory management, no GPU required. Production-focused infrastructure layer.
DeepSeek permanently reduces V4 Pro pricing by 75% to $0.435/M input tokens. Market shift toward cheaper LLM inference.
Guide on understanding and analyzing code generated by AI systems.
NotebookLM extension adds search, folder organization, bulk import, and saved prompt features.
arXiv: Reinforcement learning method for industrial scheduling that bridges simulation-to-real gap via execution semantics framework.
Piper: LLM-driven DevOps copilot where the model proposes actions from a fixed catalog, validated deterministically before human approval, never executing directly.
Crabbox.sh Pond provides runtime infrastructure for AI agents and CI/CD pipelines.
System combining human input, data agents, and AI to build self-calibrating POI map. LLM application for geospatial data.
Technique achieving 3k tokens/sec LLM inference throughput on consumer GPUs. Performance optimization for LLM inference.
NotebookLM extension adds search, folder organization, bulk import, and saved prompts for research organization and quiz/flashcard generation.
Discussion of Andrej Karpathy's Eureka Labs startup status after he joined Anthropic. Project status inquiry.
AgentKeeper v1.1: Open-source infrastructure enabling long-lived AI agents to maintain state across model switches and restarts.
Research shows AI agents consume 35-50% more tokens on unhealthy codebases across C++, Java, Python.
Article on using Jupyter notebooks for AI data analysis conversations and building AI analyst tools.