Show HN: AgentHarness – Open-source deep-research benchmark harness (Apache 2.0)
AgentHarness is open-source evaluation framework for testing AI research agents on deep-research benchmarks using ReAct setup.
AgentHarness is open-source evaluation framework for testing AI research agents on deep-research benchmarks using ReAct setup.
Open-source analytics engine inspired by Anthropic's internal tool, available on GitHub.
AI system designs real 7nm GPU in Verilog and GDSII from specifications.
Agent-ready tool extracting AI-generated fonts from images to TTF/SVG formats.
Google GKE gateway accelerates LLM inference responses by up to 92%.
Orchestration framework for deterministic execution of AI coding agents.
Plastron is a single-file HTML spreadsheet application that grows into interactive apps with formula-based cells.
Security firewall for AI agents in production. Controls access to systems, requires LLM review of destructive actions, logs activity.
Opinion piece on AI agents as workplace collaborators.
Stealth Chromium fork with fingerprint hardening and Playwright integration for Python/Node browser automation.
CLI tool providing commerce infrastructure for AI agents to conduct transactions.
User discusses challenges tracking micro-changes and reverting commits when using Claude Code and Cursor AI coding agents.
TypeScript guardrail implementation for managing AI agent cost failures locally.
Chris Lattner analyzes why GPU programming alternatives to CUDA (OpenCL, SYCL, OneAPI) failed to gain traction in AI, despite aiming to democratize AI compute.
Analysis of failed GPU programming alternatives (OpenCL, SYCL, OneAPI) and lessons for future AI compute portability.
Local-first application indexing Claude/Codex conversations into SQLite for history, search, and analytics with resumable sessions.
ChromiumFish is open-source fingerprint-hardened Chromium fork for web scraping with Python and JS SDKs.
Open-source agentic end-to-end browser testing framework executable from terminal.
Open-source plugin generating single-file HTML decks for use by coding agents.
Tool for converting AI-generated Markdown or HTML into shareable links without signup.
Research comparing LLM-based hyperparameter optimization against classical algorithms.
Open-source canvas IDE designed for agentic coding workflows, Show HN submission.
AI solved an Erdős mathematical problem; experts discuss need for safeguards.
Y Combinator AI Stack offering $20k+ in cloud credits and AI dev tool access for students.
Systems programmer guide to LLM inference covering model optimization, quantization, and practical implementation using Qwen35B model.
Tutorial on building AI agents from scratch with long-task planning capabilities and tool integration.
Gemma 4 12B multimodal model for laptops. Mobile-efficient agentic intelligence combining encoder-free design with advanced reasoning.
Google DeepMind launches 3-month accelerator for 15 European robotics startups with AI mentorship and model access.
Essay on using LLMs as reference materials for programming, discussing limitations when users lack domain expertise.
Discussion about installing third-party AI agent skills/extensions. Community question without technical depth or data.
RubyLLM 1.16 adds concurrent tool execution, Rails instrumentation, and provider proxies for LLM agent workflows. Active development update.
fftext is a CPU-based CLI tool for text summarization, fact-checking, and ELI5 explanations without GPU or cloud dependency. Open-source LLM application.
Empirical measurement of AI-written code quality compared to human code, addressing common claims with actual data from merged code.
Headline only. Malware that deceives AI security agents. Security research concept without details.
Headline only. Open dataset for distilling GPU audio2face models to CPU. Model compression/distillation research incomplete.
Open source tool to evaluate LLM performance across different languages via CLI and website.
Experimental coding agent for small (<10B) and tiny (<1B) local LLMs. Opinionated stack: Rust backend, TypeScript frontend. WIP with promising early results.
Press release for law enforcement AI platform using multi-agent framework for investigative analysis.
Open source prompt engineering framework to ensure AI coding agents remain accountable to original goals.
TokenTamer is a middleware proxy that compresses code context for LLM coding agents using AST parsing, reducing API costs by 50-80%. Alpha-stage open project.
Conceptual exploration of a minimal learning algorithm using one byte of memory to predict binary event streams. Educational content on ML fundamentals.
SuperTree is an open-source Python package for interactive decision tree visualization in Jupyter notebooks. Supports scikit-learn, XGBoost, LightGBM, and ONNX models.
Technique for cost-efficient LLM classification using lightweight models to route easy vs hard inputs.
Knowcast generates explanatory videos by analyzing concepts, creating storyboards, and rendering images into video format.
Open source framework adding continuity and multi-persona support to Claude Code LLM harness.
Meltdown: Python/Tkinter-based LLM platform with custom widgets, multi-model sessions, and tool extensions like web search and persistent memory.
CalmSEO: MCP endpoint exposing Google Search Console data, live SERPs, keyword volumes to Claude, ChatGPT, and other MCP-compatible agents.
HeadlessTracker: MCP server for crypto portfolio aggregation across exchanges and wallets, autonomous development by AI agent Hex.
Discussion of agent swarm orchestration techniques: evolution from simple prompting to centralized database, MCP protocols, and Python orchestrators with UI.
Agent-first authentication framework treating AI agents as first-class users with durable, identifiable, delegable, revocable identities for developer work.