Betlang: A tiny (50kb) programming language detection model
Betlang is a 50KB machine learning model for CPU-based programming language detection supporting 30+ languages with ranked probability outputs.
Betlang is a 50KB machine learning model for CPU-based programming language detection supporting 30+ languages with ranked probability outputs.
Microsoft announces Agent 365 for autonomous enterprise governance by 2026. Product announcement for AI agents in enterprise.
Study finds ChatGPT and AI bots made factual errors during Scottish election. Research on LLM accuracy in political context.
MulmoClaude is an open-source AI-native application platform where Claude composes tools and GUIs as plugins in a single registry, with examples including accounting systems and obligation engines.
Anthropic and Gates Foundation commit $200M for AI in health and education. Partnership announcement for public goods.
LLM application that generates cloud infrastructure designs from descriptions/sketches, producing validated code, security grades, and cost estimates.
Hocuspocus v4 released under MIT license, a WebSocket backend for real-time collaboration built on Yjs CRDT library, running on Bun, Deno, Cloudflare Workers.
macOS workspace manager for Claude Code and other AI dev agents with resumable sessions, terminals, browsers, and task management.
Framework for AI agents to design and run tests for distributed systems, generating test plans and findings reports with multi-model support.
Browser-based calculator comparing LLM API pricing across models using familiar content examples (novels, emails, contracts).
RTMX is a CLI tool for requirements traceability that integrates with AI agents via MCP, allowing agents to build against defined specifications tracked in git.
OCL Nexus Local is an open-source compute fabric for local-first AI agent development, featuring isolated Ubuntu sandboxes, Model Context Protocol support, and Docker Compose deployment.
Educational tool explaining token economics for AI APIs, comparing tokenization and pricing across Claude, Gemini, and ChatGPT.
Security vulnerability (CVE-2026-45829) in ChromaDB vector database allowing unauthenticated code execution on exposed servers.
Essay on coordination bottlenecks in AI-driven engineering teams, discussing communication failures and documentation practices.
Open-source MCP server providing safety guardrails for AI coding agents, blocking destructive operations across SQL, git, filesystem, cloud, and kubernetes.
LLM INQUISITOR is a methodology and tool for evaluating AI systems in real-world workflows, measuring stability, reliability, and safety beyond benchmark performance.
Framework for unified dev environments supporting humans, CI systems, and AI agents, addressing the multi-audience complexity of modern development workflows.
Agyn is an open-source Kubernetes platform for deploying AI agents to enterprise infrastructure with built-in security, budget controls, and secret management.
HTML-anything is an agentic HTML editor where local AI agents autonomously write and generate HTML code.
Benchmark comparison of AI agent performance across five TypeScript backend frameworks.
Toto 2.0 is a family of open-weights time series foundation models (4m-2.5B parameters) demonstrating scaling improvements on Hugging Face.
Google Search evolving into AI agent that answers, follows up, and takes actions beyond ranking links. Impact on web economics.
PostgreSQL-compatible database engine in Rust supporting relational, graph, and vector queries. Open source.
Critical analysis of LLM-generated content polluting peer review, research, and academia, causing epistemic erosion.
Google shifts focus from chatbots to agents with Gemini 3.5 Flash, emphasizing coding pipelines, autonomous research, and multi-agent coordination.
Yugabyte's Meko addresses multi-agent AI production issues with shared memory and coordination layer for state synchronization.
DNS records analysis of 39 AI company domains revealing vendor relationships, email spoofing risks, and DMARC/SPF gaps.
Open-source agentic QA harness with memory capabilities, includes live demos for testing.
Google's Gemini 3.5 Flash now available on GitHub Copilot with near-Pro quality coding at Flash-tier speed, strong tool use for agentic workflows.
BBC investigation reveals methods for manipulating AI chatbots to spread misinformation. Google and others developing countermeasures.
Professor Goose is a Socratic AI tutor using the Feynman technique. Creator fixed the system to better measure understanding through a visualization bar.
Engineer used AI coding agents to build 100K lines of production Rust code for multi-Paxos consensus engine in 4 weeks, documenting AI-assisted development at scale.
Artist without CS background proposes LLM architecture hypothesis based on biological multi-agent intelligence, testing via agent infrastructure for art business.
Cerebras running Kimi K2.6 trillion-parameter model for enterprise inference, with benchmarks on open-weight models and agentic coding systems.
Article argues legacy SEO metrics (domain age, backlinks, authority scores) are ineffective predictors in era of AI agent web usage.
Auto Agent Protocol (AAP) defines A2A specification for AI agents to buy cars, including inventory and pricing APIs.
Framework describing three generations of AI applications: conversational (chat), delegative (task automation), and collaborative (agent partnerships).
Guide to managing fleets of AI agents by persisting outputs to hierarchically organized markdown artifacts for durability and tracking.
Personal account building AI Chief of Staff agent to handle email, Slack, task coordination, and calendar management for managers.
Claude skill bundle for security testing with 574 vulnerability patterns, 51 skills, 15 commands, and Burp MCP integration for red-team and bug-hunting work.
Google announces Gemini Omni Flash, multimodal model for video and content generation.
Slax Reader CLI library enabling AI agents to access read-later services for content retrieval.
SRM framework for detecting slow-burn risks in AI agent sessions before execution.
Chrome extension enabling terminal coding agent 'zot' to operate browser tabs through browser_action tool.
Nucleus: permissions and policy enforcement framework for AI agents with enforcement guarantees.
Mirage: virtualization layer enabling AI agents to access multiple backends via unified virtual filesystem and bash commands.
ExecuTorch MLX Delegate enables GPU-accelerated PyTorch inference on Apple Silicon Macs via MLX framework.
Technical approach for scheduling data egress during peak solar production to reduce grid impact of AI training.
OEP attack: poisoning self-evolving LLM agents via locally correct non-transferable experiences. Security vulnerability in memory-augmented agents.