Open-source DCF engine based on Damodaran's datasets with LLM narratives
Local-first DCF valuation tool using LLM narratives on top of financial calculations, educational project.
Local-first DCF valuation tool using LLM narratives on top of financial calculations, educational project.
Miguel is an AI agent that modifies its own source code, self-improves capabilities, sandboxed in Docker with validation.
Rampart is an open-source security firewall for AI coding agents with 40+ rules blocking credential theft and exfiltration.
Article on running 70B LLMs on Nvidia RTX 5090 with FP4 quantization benchmarks, member-only content.
Open-source vision-first browser agent that automates web interactions using visual understanding instead of DOM selectors, reducing token waste and script fragility.
Draxl is a source code format with stable AST node IDs for agent-native code editing at scale.
Artifice is a multiplayer strategy game for AI agents with diplomacy and fog of war, open source implementation.
Web interface for Claude Code featuring real-time visualization of model steps/tool calls, chat UI, and session management.
AI agent integrated into Appium mobile test automation platform. Analyzes live device screens and generates selectors/XPath code in multiple languages for test automation.
Web tool that analyzes and enriches rough AI prompts into structured, optimized versions across five dimensions.
OverflowML tool auto-detects hardware and applies optimal memory strategies to run AI models larger than GPU VRAM. Supports NVIDIA, Apple Silicon, AMD, CPU.
Technique to top HuggingFace Open LLM Leaderboard without training or weight merging, using prompt engineering and evaluation manipulation.
StrongDM open-sourced attractor: natural language specs and implementations for unified LLM client, coding agent loop, and DOT-based pipeline runner in multiple languages.
Summary of Pragmatic Summit talk about Uber's use of AI in development. Limited technical details provided in excerpt.
Tool to sync configuration between Claude Code and Codex, automating shared parts while flagging manual migration tasks.
Brief report that Claude Code causes 90% slowdown when used with local LLMs. Minimal details provided.
Security researcher demonstrates GPT-4 training data leakage exposing OpenAI's EPHEMERAL_KEY through repeated bypass attempts with 75% leak rate.
Control plane/policy engine for AI agent actions with human approval queue and deterministic YAML policies. Self-hosted tool for production agent safety.
AlphaEvolve-inspired agent using iterative code generation and scoring for Pokemon task. Demonstrates LLM agents writing and improving code autonomously.
AI agent that learns from execution errors to refine its own decision rules. Demonstrates adaptive agent behavior and self-improvement.
TCP proxy preventing AI agents from executing destructive database operations. Developer tool for agent safety and authorization control.
Open Prompt Hub platform enables sharing prompts instead of code so AI agents can generate customized software from prompt specifications.
g0 is a unified security control layer for AI agents with static/behavioral analysis, 1,180 rules across 12 domains, supporting 10 frameworks.
Modulus desktop app enables multiple coding agents with shared cross-repository project memory to understand dependencies across separate codebases.
Federal judge blocked Perplexity's Comet AI shopping agent from accessing Amazon after lawsuit alleging concealment and unauthorized web scraping.
MemoTrader marketplace for AI-human messaging offers MCP server enabling Claude agents to register, fund, and contact humans with minimal configuration.
rolvsparse compute primitive benchmarks matrix arithmetic optimization achieving up to 82x speedup on DeepSeek-R1 and Llama 4 models.
Free tool that scans system prompts against 12 attack categories to identify prompt injection vulnerabilities with example exploits and fixes.
TokenZip Protocol proposes passing pointer references between LLMs instead of full token sequences to reduce context usage.
Open-source runtime for Claude Code that adds security guardrails between AI agents and shell execution. Enables safer autonomous agent operation.
Claude Code skills pack providing 12 terminal commands for startup founders addressing strategy, market fit, and business validation.
Benchmark showing token optimization for AI coding agents isn't straightforward. Pre-indexed context via MCP reduced costs 24% despite 20% token increase.
Tauri desktop app for orchestrating multiple Codex agents across local workspaces with project management and conversation interface.
Python SDK detecting and tokenizing PII on-device before LLM processing. Enables safe handling of sensitive data in AI agent pipelines.
Open-source Node.js framework for building programmatic AI agents. Agents adapt execution based on instructions, tools, and memory instead of static workflows.
NBER working paper modeling how generative and agentic AI shapes human learning incentives and information ecosystem evolution.
Technical specification design for APIs serving AI agents instead of applications. Compares Skills, Tools, and MCP standards for agent tool calling.
Article on AI's dual impact for open-source: Claude helping find bugs in Firefox while raising concerns about training data usage.
AI agents trained on 1M+ lines of F* and Pulse code/proofs to build provably correct implementations of classic algorithms and data structures.
Autoautoresearch extends Karpathy's hyperparameter search with LLM agents to address blank page problem in AI-driven research.
Best practices for hosting and authenticating remote MCP (Model Context Protocol) servers. Developer guide for agent infrastructure.
Research showing AI agents perform worse with 100k tools vs fewer tools. Challenges tool scaling assumptions in agentic systems.
Google releases Gemini multimodal embeddings supporting video and PDF inputs. Enables richer semantic search across media types.
Technical write-up on SQLite concurrency patterns in Go while building a desktop AI IDE. Developer tools and architecture lessons.
Legal case blocking Perplexity's AI agent from autonomous Amazon shopping. Early test of agentic commerce regulation.
Methodological critique of 'First Proof' paper evaluating AI capabilities on research-level math problems. Identifies experimental design flaws.
Open-source MetalRT inference engine for Apple Silicon outperforming llama.cpp and MLX. Includes RCLI voice AI pipeline; mic-to-response entirely on-device.
One-command deployment tool for AI-generated code from Claude Code or Cursor. No Docker/YAML; supports Mac and Linux with auto-runtime detection.
Tutorial series on training GPT-2 from scratch investigating learning rate hyperparameter choices for improved test loss optimization.
Open-source SEO/AEO tool tracking AI agent citations and visibility in AI-powered search. Helps merchants prepare for agent-driven commerce.