Show HN: Inferock-bench – per-call billing receipts for OpenAI and Anthropic
Inferock-bench: independent per-call billing receipts and failure tracking for OpenAI and Anthropic API usage.
Inferock-bench: independent per-call billing receipts and failure tracking for OpenAI and Anthropic API usage.
Entire is a distributed Git system designed for AI agents to manage code and collaborate.
AgentPeek is self-hosted browser control panel for Claude Code sessions that persists across disconnections and browser reloads.
Terminal Control tool enables AI agents to control and test terminal applications via pseudo-terminal instead of text output.
Green is a Clojure/babashka library for idempotent DevOps CLIs with EDN workflows, OpenTofu, and Ansible integration.
Red is a TypeScript/Bun library for idempotent DevOps CLIs using YAML workflows, OpenTofu, and Ansible provisioning.
Open source GitHub project adding explain-like-im-five prompting rule for Claude to reduce AI output fatigue.
Tool integrating Google Calendar with AI agents for task scheduling and automation.
Decision tree framework for selecting appropriate memory strategies for AI agents, addressing common memory design pitfalls.
Open protocol enabling any website to work with AI agents through standardized interfaces, addressing agent compatibility issues.
Measurement of token costs for AI agents fetching web content, showing Wikipedia costs 68k tokens raw vs 950 tokens summarized, identifies limitations with JavaScript-rendered sites.
Python proxy that bridges Atuin Hub AI to OpenAI-compatible endpoints, enabling use of multiple LLM backends with Atuin shell history tool.
Review comparing AI research tools including Claude Science, evaluating their capabilities for scientific research workflows.
Hana JIT is an LLVM-backed Python JIT compiler that compiles Python functions and NumPy code to native machine code with a genetic-algorithm superoptimizer for performance.
Curated repository of step-by-step guides for building technologies from scratch. Educational resource maintained by CodeCrafters.
Ethereum deployed AI agents for security auditing and discovered libp2p vulnerability, demonstrating practical AI agent application in bug hunting.
Open-source Chrome extension for session recording: captures network calls, console logs, screenshots locally without cloud backend.
Agency runs 10 AI agents on launchd timers without frameworks, showing minimal-dependency approach with guardrails for safe autonomous operations.
Demonstrates running a full AI agent-powered company operation via web scraping and live charting, showing practical autonomous business system.
Insurance agency built MCP server for AI agents to generate disability insurance quotes using Claude Code, demonstrating practical agent deployment.
Essay comparing China's open-source AI models vs America's closed proprietary approaches and implications for global adoption.
Runko: self-hosted monorepo platform on Git with stacked review, path ownership, and AI agent support. Apache-2.0 licensed.
Stub article on agentic test processes, LLM benchmarks, and agentic coding observations. No content provided.
Guide on governance mechanisms for autonomous AI agents, addressing containment strategies to prevent unauthorized actions.
Empirical measurement of frontier coding agents showing they misrepresent their work completion and accuracy despite claims of capability.
Comparison of GPT-5.6, Grok 4.5, Claude, and Meta's Muse Spark building identical apps with cost/latency benchmarks.
Case study migrating production AI agent from Claude Opus to GPT-5.6 Sol, with benchmarking and migration guide.
Technical walkthrough of GPT-2 implementation in JAX, explaining parameter counts from embedding dimensions and model architecture components.
Coder_eval: benchmarking framework for evaluating AI coding agents (Claude, Codex, Gemini) with YAML tasks, scoring, and CI integration.
Testing framework providing unit test-style validation for agent skills and capabilities.
Question about potential pricing discrepancies for GPT-5.6 models on OpenRouter platform.
Multi-pass agent workflow generated 50k-word novella with planning, drafting, review, and quality checks across 77 markdown files in 8 hours.
Cactus v2 on-device inference platform with cloud fallback, confidence-based routing, PyTorch converter, 4-bit quantization, and GPU acceleration.
Bunrun: agent-configured dashboard for managing local dev environments. LLM-assisted YAML configuration for development tooling.
Stripe Radar fraud detection system for AI agents claims 86% Sybil fraud reduction.
Terminal text editor/IDE with split-panel diff review, LSP support, and SSH remote editing. Developer tool with technical depth and original implementation.
VS Code/Cursor plugin enabling local LLMs as coding agents with automatic model profiling and tunnel setup for remote connection.
Case study on performance degradation after upgrading an AI agent system.
PDF research on AI-generated papers passing peer review undetected at ACL conference.
Research claim about GPT-5.6 Sol Ultra producing proof of Cycle Double Cover Conjecture. Mathematical/research application of LLM with high engagement.
Open-source AI voice generation tool. Limited details provided.
AI application providing feasibility assessments for business ideas.
CUDA-native AI guardrail kernel written in C++20, 100% branchless implementation for LLM safety.
Open-source handbook on agent harnesses from Tencent, covering auditability and editability of coding agents. Original documentation on AI agent infrastructure.
9lives is a self-healing test runner that automatically fixes broken Playwright tests by classifying failures and healing them in tiers, designed to prevent AI agents from unnecessarily rewriting tests.
YC S24 company shipping Scribe (style-learning dictation) and free AI dictation using fine-tuned Llama 3.1 8b.
Technical retrospective on generative spatial AI advances: text-to-mesh, video, 3D models, world generators, CAD integration.
Title only. Discusses LLM failure modes involving delusional/circular reasoning in chatbots.
Show HN: GenUI enables AI agents to generate real, interactive SwiftUI interfaces for iOS/macOS instead of text output.
Research on memory requirements for LLMs beyond weight storage.