Open-source AI battle arena, plug your coding agent in and it fights bots
Open-source AI battle arena platform for testing and competing coding agents with live tracking and bot API.
Open-source AI battle arena platform for testing and competing coding agents with live tracking and bot API.
Open-source on-device AI apps running locally on NPU hardware without cloud dependency or data transmission.
Research paper on using LLM-based agents with micro-profiling tools to optimize CUDA GPU kernels, treating profilers as expert surrogates for kernel optimization tasks.
Local-LLM-only AI content pipeline publishing 100+ autonomous posts with human review, no cloud dependency.
AgentsProof: Open-source project for testing and evaluating AI agents with run tracing, test cases, and shareable reports.
Experiment testing local LLM behavior when given dangerous production database deletion capabilities and constraints.
Phlox-GW: Open-source self-hosted LLM gateway with SSO, cost accounting, guardrails, audit logs, and clustering features.
Survey of 696 SREs shows only 8% have AIOps in production; 73% cite lack of trust as barrier to AI agent adoption.
Microsoft CEO Satya Nadella warns enterprises against sharing IP with frontier AI labs to avoid reverse information attacks.
Implementation of neural network in SQL using Xarray-SQL array database library. Demonstrates mapping n-dimensional arrays to 2D tabular representation.
PlanWright is an MCP-driven control plane for AI coding agents, enabling planning, implementation, and triage with full documentation.
Mindshub is an open-source AI coworker alternative to Claude, emphasizing inspectable, modifiable, and forkable AI systems.
Educational content about understanding LLM parameter counts and model scaling. Machine learning research.
OpenAI released GPT-5.6 models that cleared high cybersecurity capability threshold, automating vulnerability discovery and attack scenarios.
Anthropic interpretability research exploring whether vision-language models think in text tokens, testing Jacobian-based methods on VLMs like Qwen2.5-VL.
Production case study migrating 100B tokens/week from GPT-5.3 to open-weight MiniMax M3, achieving 55% cost reduction for QA agents.
Promptster: AI fluency platform measuring engineering teams' proficiency with Claude Code and other AI coding tools.
Using local LLMs to block unwanted website content. Browser-based LLM application. Limited details.
Microsoft CEO proposes new patent framework to address IP theft concerns from AI models trained on corporate data.
Discussion of agentic science and autonomous labs discovering new compounds, exploring AI's role in scientific discovery.
Durable execution framework for AI agents using object storage and job queues. Solves agent failure recovery.
Benchmark comparing 5 LLM models across 13 Ruby codebases, measuring code generation vs. navigation capabilities.
Demo of AI-powered Big Mouth Billy Bass toy using Strands agent framework for autonomous control.
Research on decision-making processes in AI shopping agents, analyzing product selection behavior across models.
Gimlet uses formal verification to validate AI-generated GPU kernels, improving trust in agent-optimized code.
CLI tool providing Claude/GPT coding agents direct access to financial data: stock prices, options, SEC filings via structured API.
Guide on designing and deploying AI agents for enterprise business team workflows and adoption.
Shared memory/context management system for AI tools and teams. Infrastructure for LLM applications.
Show HN post about a backend for AI-generated apps. Minimal details provided.
Platform crowdsourcing adversarial attacks and jailbreaks against AI agents, with rewards for vulnerability submissions.
Technical case study: optimization techniques reducing token consumption in AI agent systems by 94%.
AutoFlow automates user guide generation from screen recordings using chrome extension and ML transcription. LLM-powered product documentation tool.
Security framework for AI agent accountability using identity verification and access controls for privileged operations.
Show HN tool for tracking cloud and AI token spending. Helps monitor LLM usage costs.
Docs.dev platform for AI-assisted documentation where Claude Code agents draft content and users edit in place, with keyboard.dev integration.
Python/FastAPI middleware for defending LLMs against prompt injection attacks. Security developer tool.
Machine learning research on learning from experience vs. curated datasets. Training methodology comparison.
Prompt engineering technique to generate property-based tests using LLMs. Developer tool for testing.
Cryptocurrency-based anonymous proxy for LLM API access. LLM infrastructure. Minimal details.
Browser extension adding folders, search, bookmarks to Claude.ai conversations. Local-only, no data collection.
Benchmarking of Apple's SpeechAnalyzer against Whisper on 5,559 utterances. SpeechAnalyzer most accurate, 3x faster than Whisper Small. Transparent results published.
Docker-based template for AI development with multi-provider CLI support and isolation. Works on Linux, macOS, WSL2.
Article about using LLMs with real-time market data for stock research. Generic LLM application overview.
A/B testing results for forms.app's AI generator hero section achieving 2.6x signup increase. Includes LLM-powered form generation feature.
Domain-specific language for data pipelines that reduces LLM token usage by 33%. LLM optimization.
WordPress agentic system for debugging and development that can inspect actual site configuration and errors.
Discussion of AI safety concerns as LLM-based agents become capable of unsupervised work over extended periods.
Comprehensive guide to 900+ developer tools for AI coding agents, accessible via CLI and MCP integration.
Show HN: Yes-Brainer tool enabling multiple LLMs to debate in browser with bring-your-own-key. Multi-agent LLM application.
Harpist tool converts undocumented websites into refined APIs using code analysis and HAR file recording.