We Built 115 Free SLMs for Specific Agentic Tasks
Release of 115 free specialized small language models designed for specific agentic task applications.
Release of 115 free specialized small language models designed for specific agentic task applications.
Research on measuring AI manipulation and its potential to alter human thought and behavior through deceptive interactions with language models.
Analysis of cost inefficiencies in AI agent systems: infinite loops, redundant tool calls, and hallucinations causing 40% budget waste.
Observability platform integrating with LLM-based coding agents (Claude, Codex, Cursor) for plain-language app interrogation and alerting.
Cryptographic delegation system replacing API keys for secure AI agent authorization and permission management.
HarmActionBench research: GPT and Claude agents fail safety checks when instructed to perform harmful actions with tools.
Open-source security testing framework based on OWASP standards for evaluating AI models and agent vulnerabilities.
TeamMind adds persistent memory layer to Claude Code, runs locally without API key.
dbt-skillz tool compiles dbt projects into Claude Code skills to improve coding agent performance on data tasks.
Part 32g of LLM training tutorial series covering weight tying intervention technique.
Google's AI model compression research paper available on arXiv since April 2025.
Claude plugin enabling coding agents to follow software design principles and best practices.
Local proxy tool enforcing guardrails for AI agents using HTTP x402 payment standard.
Llumen is a lightweight LLM chat application.
Security research on AI coding agents running on developer machines without visibility; Sysdig TRT building detection layer for agent behavior.
Open-source protocol enabling AI agents to discover services, negotiate terms, and settle payments via encrypted channels.
Technical article part 2 about SPy language semantics implementation.
Personal reflection on Leon AI 2.0 open-source assistant development; philosophical stance against hype-driven development.
Technical lessons on building AI data analyst agents; infrastructure insights on optimizing agents for data workflows.
Guide comparing LLM frameworks available in 2026 for developers.
Local alternative to cloud LLM APIs. Stack for running domain-specific models on commodity hardware without external providers. Open source project.
Claude Auto Mode allows AI agents to make decisions about safety and task execution. New capability for autonomous agent behavior.
AI tool for validating startup ideas through stress-testing; LLM application.
Categorizes agentic AI tools across 11 categories. Limited technical depth; appears to be taxonomy/overview.
Vectree generates interactive SVG visualizations using LLMs to explain complex concepts. Educational application with visual learning focus.
Approva: Open core human approval infrastructure for AI actions. Governance layer for autonomous AI systems.
Comparative safety testing results across 6 LLMs (GPT-4o, Claude, Grok, DeepSeek, Gemini) with 3,360 test cases.
Research on trade-off between expert personas improving LLM alignment while reducing factual accuracy.
ARK runtime reduces AI agent context overhead by 99% through dynamic tool schema learning. Persists decisions across runs for improved efficiency.
Red team security testing against AI agents with production access. Four social engineering attacks tested; agent resistance evaluated.
Local-first Python agentic backlog generator using Ollama. Generates epics, features, acceptance criteria. No API keys, demonstrates agentic patterns.
Cryptographic system for autonomous AI agents using Schnorr signatures and zero-knowledge proofs for trust verification without API keys.
Open-source tool using AI to automatically rename PDF files based on content.
AskAlf orchestrates teams of specialized AI worker agents for specific domains, automatically configuring and managing them for 24/7 operation.
Technical guide on using Git worktrees and SQLite for coordinating multiple parallel AI agents working on the same monorepo without infrastructure overhead.
LisPy is a Scheme-like Lisp interpreter for AI agent orchestration that represents agent state as executable s-expressions instead of inert JSON.
Researcher seeking production engineering partners to advance a 7-year deterministic AI substrate project as alternative to probabilistic token-based models.
Experiment testing three AI agents' willingness to push back on feature requests, finding models trained to be helpful tend to over-complicate solutions instead of refusing.
Permission vulnerability in Claude Code where first-token-only flaw allows dangerous actions. Triage bot dismissed as informational.
Analysis of hidden costs in AI agent usage: while token prices drop, multi-agent reasoning loops and self-critique increase bills despite cheaper per-token rates.
FlowScript is a typed reasoning system for AI agent memory that addresses limitations of flat fact storage and vectorized chains by enabling contradiction tracking for better agent reasoning.
News aggregator covering AI agents and agentic AI developments for builders and founders.
Research taxonomy documenting 50 failure modes in multi-agent AI systems. Original research for reliability improvement.
Open social network platform for AI agents to interact and coordinate. Enables agent-to-agent communication.
Discusses cost accounting for AI agent customer support ticket resolution.
Ente releases Ensu, offline LLM app emphasizing local models for privacy and control. First release with multi-device support.
Validating compiler for agent and human-in-the-loop workflows. Generates typed code from workflow definitions in Go, Python, Rust.
Drift detects architectural degradation in codebases using AI-assisted coding. Developer tool for maintaining code quality.
Open-source tool that connects AI agents to 40+ services (Gmail, GitHub, Notion) via CLI authentication and skill generation.
MCP server integrating Transloadit media pipelines with AI agents for video, image, PDF processing workflows.