Skillwave is an autonomous agent orchestrator that decomposes goals into tasks, creates subagents with distinct roles, and executes via async communication loop until completion.
Career advice article about knowledge transfer and team transitions.
Customermates open-source, self-hostable CRM with AI-first design.
Chardet character encoding library rewritten from scratch using Claude. Detailed technical account with conversation transcripts showing AI-assisted development process.
AgentLair credential vault for AI agents preventing environment variable exfiltration in supply chain attacks.
Discussion thread on limitations of current LLM code generation models across different complexity scenarios.
Explores credential management approaches for AI agents requiring secure access to passwords and authentication systems.
Google DeepMind research study on AI's persuasion and manipulation capabilities.
Zinc LLM inference engine in Zig enabling 35B model inference on consumer AMD GPUs via Vulkan.
API design principles for LLM consumers. Reducing Claude's healthcare API calls from 72 to 8 through agent-focused redesign.
Analysis of AI-generated patches passing CI tests but failing production. 20% breakage rate in vulnerability fixes.
Open source email infrastructure for AI agents. Send, receive, search, extract codes. Deploy on Cloudflare. Integrates with Claude Code and AI agent platforms.
Open source library of 450+ modular agent skills for medical research. Works with OpenClaw, Claude with scientific integrity constraints.
Open source macOS terminal multiplexer for running AI agents in parallel with notifications. Built for agent workflows.
Analysis of accelerating AI tool/framework releases tracked via HN, GitHub, npm, PyPI showing ecosystem growth rate.
Analysis of how AI agents integrate third-party tools into code generation and product decision workflows.
Empirical study showing verification steps degraded AI agent performance across 29 tests. Original experimental research.
APIEval-20 benchmark dataset for evaluating black-box API test suite generation using LLMs and schemas.
GPU profiling tool that diagnoses performance bottlenecks beyond utilization metrics. Minimal details provided but relevant tool.
MCP server for AI agents to select appropriate cloud services with current pricing and compatibility data. 74 services, no API key required.
Stanford research showing AI vision models generate images not in training data through hallucination mechanisms.
TRIBE v2: Predictive AI model of human brain responses to visual, auditory, and language stimuli from neuroscience research.
R package that converts Excel workbooks to standalone R scripts with formula recreation and verification against cached values.
LLMnesia: Local-first search tool for AI conversation history across ChatGPT, Claude, Gemini, and other platforms.
Analysis of Meta's legal losses and liability implications from internal social science research on platform effects.
WhisperFlow: Free, open-source speech-to-text tool for macOS. On-device processing, no cloud upload, no account required.
Lightweight desktop tool for quick text capture and organization across applications, useful for knowledge management workflows.
Google's internal AI tool 'Agent Smith' automates coding tasks. Became so popular access was restricted. Limited technical details.
Official codebase for paper on AI agents for embedded/IoT systems development. Addresses hardware-in-the-loop constraints and physical behavior coupling.
Local tool that analyzes Claude Code session logs to identify token usage patterns and wastage, processing files without cloud connectivity.
BeSafe-Bench evaluates behavioral safety risks of multimodal agents performing complex tasks in functional environments with high-fidelity evaluation framework.
AutoB2G uses LLM-driven agents to automate building-grid co-simulation workflows and reinforcement learning policy development for building cluster control.
Framework for knowledge engineering and process mapping in airport operations using machine-readable representations to resolve data silos and semantic inconsistencies.
GUIDE mitigates domain bias in GUI agents using real-time web video retrieval and annotation to improve planning and grounding for domain-specific software operations.
AIRA_2 addresses bottlenecks in AI research agents: synchronous execution, generalization gaps in validation, and limited LLM operator capabilities through architectural improvements.
CADSmith generates CadQuery code from natural language using multi-agent pipeline with iterative refinement loops for geometric validation and error correction.
PAPO (Process-Aware Policy Optimization) integrates process-level evaluation into GRPO via decoupled advantage normalization to improve reward model signals in reasoning tasks.
DesignWeaver research on text-to-image generation for product design; study with 12 experienced designers showing how experts explore design spaces; enables novices to create professional visuals.
Sommelier framework enables scalable multi-turn audio preprocessing for full-duplex Speech Language Models supporting real-time conversational interaction.
A-SelecT automatically selects optimal timesteps in Diffusion Transformers for discriminative representation learning, improving efficiency of diffusion-based pre-training.
Empirical study measuring behavioral consistency of LLM-based agents (Claude, GPT-5, Llama) on SWE-bench software engineering tasks, revealing how variance impacts reliability.
ETA-VLA optimizes Vision-Language-Action models for autonomous driving by reducing computational burden of multi-view temporal reasoning through token adaptation and LLM sparsification.
Data-centric study proposing strong supervision framework for audio pre-training models, drawing from vision pre-training approaches to improve representation learning.
UCAgent end-to-end agent for automated IC block-level functional verification addressing bottlenecks in semiconductor design verification.
IncreRTL framework for incremental RTL generation adapting to evolving requirements using requirement-code traceability links.
ReCUBE benchmark to evaluate how effectively LLMs utilize repository-level context during code generation on large codebases.
Survey of reinforcement learning applications in infectious disease control and epidemic response optimization.
Theoretical framework for learning causal representations from limited environments with finite-sample guarantees bridging causal models and latent factor models.
MAGNET system for decentralized autonomous generation and training of domain-expert language models using autoresearch and BitNet ternary training.
MedBench evaluation framework for agent-based medical AI using multi-step clinical dialogue simulation between physician and AI system.