Show HN: Evalcraft – cassette-based testing for AI agents (pytest, $0/run)
Pytest plugin using cassette-based recording for AI agent testing. Records LLM calls once, replays deterministically in CI with zero API costs.
Pytest plugin using cassette-based recording for AI agent testing. Records LLM calls once, replays deterministically in CI with zero API costs.
Modular parametric CAD system with AI-assisted design capabilities built with Python and Node.js.
Real-time global intelligence dashboard with AI-powered news aggregation, geopolitical monitoring, and multi-model summarization chain.
Security scanner testing AI agents against 150+ attack probes including prompt injection and tool misuse vulnerabilities.
CLI tool managing Model Context Protocol servers across 13 AI clients from single config. Includes 97 curated MCP server templates.
OpenClaw general availability on Amazon Lightsail: open-source self-hosted autonomous AI agent with browser integration and Bedrock model provider.
Interactive AI-powered sales chat demo for fictional luxury EV brand. Explores conversational advertising with product knowledge retrieval.
Case study: Ayrshare's dev team used multi-agent Claude work via Cursor to build scalable rate-limiting infrastructure for social media APIs.
Descript uses OpenAI models for multilingual video dubbing, optimizing translations for meaning and timing.
MUD game built with Claude Code featuring agent pipeline: game-designer agent proposes balance changes, playtest agent QA-tests gameplay.
Discussion questioning whether subagents in AI coding assistants create manager-like behavior, reducing flagship models' reasoning autonomy.
LLM-based email parser tool handling brittle format variations with schema validation, retries, and backend integration.
Conversational AI agent for Kubernetes cluster operations. Combines visual dashboard with agentic workflows to reduce context switching between monitoring and CLI tasks.
Open-source maintainers face harassment from AI-generated code contributions. matplotlib and other projects implement human review policies for AI-submitted code.
iPad agentic coding tool using Claude that autonomously reads codebases, plans changes, edits files, and pushes to GitHub. 7 integrated tools execute locally with token-streaming API.
Hopsworks platform as runtime for coding agents. Analysis of container vs wrapper approaches and MCP integration for Claude Code, Codex, Gemini.
Voice-first AI language learning application built with Next.js and Gemini APIs. Technical discussion of browser-native lip sync implementation challenges.
Tool that generates unified context files (CLAUDE.md, .cursorrules, etc.) for AI coding tools from a single codebase scan, solving maintenance overhead.
Platform for building and deploying autonomous AI agents across decentralized networks. Combines learning environment with infrastructure for cross-chain finance and trusted execution.
Safety guardrails tool intercepting dangerous commands from both humans and AI agents. Works universally across shells with blast radius detection and context analysis.
Side project: 1v1 strategy game where LLMs compete. Results show LLMs can generate functioning bots but underperform vs. human players.
Article exploring optimal human-agent collaboration in software development. Discusses whether developers should inspect all code or focus on outcomes.
Open-source helpdesk software version 7.0 adds AI features without vendor lock-in. Supports pluggable LLMs while maintaining data sovereignty.
WavSLM: single-stream speech language model using WavLM distillation for autoregressive speech generation without text supervision.
GALACTIC: counterfactual explanations method for time-series clustering identifying transitions across cluster boundaries.
FairFinGAN: WGAN-based framework for generating synthetic financial data with fairness constraints to mitigate bias.
Geometric-aware quantization technique preserving SO(3)-equivariance in GNNs for efficient molecular simulations.
InfoFlow KV method for efficient long-context RAG by recomputing KV caches with information-flow awareness to optimize prefilling.
TS-BOSS: time series extension of Best Order Score Search for causal structure discovery in multivariate temporal data.
Deep adversarial framework for human activity recognition from wearable sensors addressing inter-subject variability.
TopKGraphs method for node similarity estimation using Jaccard-biased random walks for clustering and recommendation tasks.
Sheaf Neural Networks extension with learnable restriction maps to address oversmoothing on heterophilous graphs.
Interpretable prototype parts-based neural network for medical tabular data with discretization of diagnostic norms for improved trust and transparency.
LWAIL framework for imitation learning from state-only demonstrations using adversarial Wasserstein methods in latent space.
Study of honesty elicitation and lie detection methods in LLMs using censored models as natural testbed for secret knowledge extraction.
POET-X framework for memory-efficient LLM training using orthogonal transformations to improve stability while reducing memory consumption.
Token-wise adaptive compression method for KV cache in LLMs to reduce memory footprint during inference without expensive retraining or performance loss.
Vision-language models show massive context-dependent affordance drift when primed with different agentic personas, affecting action prediction.
Empirical study showing distributed GPU training scaling fails due to network and fabric effects overlooked by training frameworks.
AMV-L: memory lifecycle management system for long-running LLM agents that bounds tail latency through intelligent retention policies.
SkillNet: open infrastructure for creating, evaluating, and sharing reusable AI skills to enable systematic knowledge accumulation across agents.
Multimodal LLMs fail via induced numerical instability when adversarial images trigger optimization of inference-time instability.
Act-Observe-Rewrite: multimodal LLM agent learns robot manipulation by synthesizing Python controller code between trials via visual feedback without gradient updates.
LLM-based approach for predicting antibody binding affinity against SARS-CoV-2 using machine learning on experimental antibody data.
AI agents using LLMs for self-monitoring exhibit self-attribution bias when evaluating their own actions, showing leniency compared to external evaluation.
Spinverse: differentiable physics method for reconstructing tissue microstructure from diffusion MRI using Bloch-Torrey simulation.
iAgentBench: benchmark for evaluating information-seeking AI agents on cross-source sensemaking and evidence integration tasks.
Method improving physics-informed neural network accuracy through last-layer retraining on PDE solutions.
Empirical study examining whether speech denoising improves zero-shot ASR with SAM-Audio and Whisper models.
Method for allocating compute resources across prefill-decode disaggregated LLM inference while meeting SLO constraints.