Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
Gaia2 benchmark framework for evaluating LLM agents in dynamic, asynchronous environments. arXiv research project.
Gaia2 benchmark framework for evaluating LLM agents in dynamic, asynchronous environments. arXiv research project.
Legal analysis of Amazon v. Perplexity case on CFAA applicability to agentic AI services. Preliminary injunction discussion.
AI sales agent tool that automates personalized follow-ups to prospects via email/SMS with prospect discovery features.
Study on stress-testing LLM safety claims. Method for detecting catastrophic failures and limitations of current AI safety approaches.
Desktop control via AI through Model Context Protocol (MCP). Integrates with Claude, Cursor, Windsurf, Zed. Local execution, no telemetry.
GitHub Copilot deprecates GPT-5.2 and GPT-5.2-Codex models across chat, edits, and agent modes.
Aquifer is an MCP runtime that coordinates HTTP traffic from distributed AI agents using durable queueing and rate limiting.
Local-first PDF reader with note-taking and keyboard navigation for studying technical books, integrates Claude/Codex.
LEAP study surveys AI experts and superforecasters on AGI timelines and AI impact predictions through 2040.
Research proposal for decoupled RISC-LLM architectures using circadian synaptic consolidation to address catastrophic forgetting and reduce latency.
Developer tool connecting code editor and REPL for unified debugging sessions using UNIX philosophy.
Kodiqa Agent: AI coding agent supporting multiple LLMs locally via Ollama or cloud APIs, with 78 slash commands, RAG, sub-agents, and plugin support.
arXiv research on tree-like self-play method for improving secure code generation in LLMs through iterative learning.
Developer orchestrated 3 LLMs to build Chrome/Firefox plugin ranking restaurant dishes in 2 hours, documenting workflow.
Analysis comparing scientific formulas and LLM compression mechanisms as analogous structural approaches to describing phenomena.
Analysis of AI-driven layoffs in 2024-2025, questioning whether efficiency gains were truly automation or cost-cutting pretexts.
Title mentions security research on Jane Street LLMs for backdoors. No content provided to evaluate.
Open source self-improving AI agent with layered memory, cron scheduling, and reusable skills system by Nous Research. Community web UI.
Research study from UT Austin showing LLM memory systems have 95% error rate in long-term fact retrieval using vector-based storage.
Meta reveals thousands of Instagram accounts hijacked through exploitation of its AI chatbot. Security incident report.
Tool to share and resume Claude Code sessions via Git branches, enabling team collaboration on AI-assisted coding work.
Tool for running and coordinating multiple Claude Code sessions reliably. Developer utility for managing AI coding agents.
Developer tool displaying Claude Code and OpenCode activity on Garmin smartwatch via BLE. Real-time monitoring of AI coding assistants.
macOS app that creates hallucinated OS with Claude-generated interactive HTML apps in sandboxed iframes. Working JavaScript computations.
Nature paper demonstrates LLMs can transmit hidden behavioral traits through model distillation, affecting downstream model training.
AgenticRL framework for autonomous robot navigation using self-refining reinforcement learning to reduce manual reward engineering.
ZML is production inference stack decoupling AI workloads from proprietary hardware, supporting NVIDIA, AMD, TPU, Trainium with optimized performance.
Jeju: local-first agent harness with inspectable execution runs for AI agents.
Anthropic reports Claude AI is scaling faster than anticipated.
Training-free method for single-image diffusion models, enabling image generation without per-image training.
Hermes Agent: open-source persistent AI agent with memory, skill creation, and multi-platform integration for personal use.
Workflow tool using AI to validate startup ideas through structured research and fact-checking before implementation.
Multi-layer policy framework for securing AI agents at runtime in enterprise environments. Technical security architecture discussion.
Experimental Rust port of React Compiler for integration with tooling ecosystems like SWC.
Guide on configuring MCP (Model Context Protocol) servers in IDEs. Security/configuration focused.
Collection of agent skills for OpenAI's /goal mode based on published best practices documentation.
Interview with OpenAI Codex Tech Lead on AI-assisted software engineering practices and tools.
AI sales agent automates product demos. Practical LLM agent application for business automation.
Browser extension providing reader-mode for web pages using local LLMs. Practical LLM application tool.
Essay on learning stalls when over-relying on AI systems without understanding internals. Educational perspective on AI limitations.
Research on productivity effects of different generations of AI coding tools, comparing writing vs. shipping outcomes.
Strategy for large companies to reduce AI costs by adding local LLM filtering layer.
Shell plugin integrating Pi coding assistant into Noctalia for IDE-like development experience.
CLI tool for imperative AI workload orchestration.
TuringLLM: open-source project using LLMs as step functions in universal Turing machines with markdown state/instruction representation.
Guide for running and optimizing LLMs, VLMs, and diffusion models on NVIDIA DGX Spark with vLLM.
Busbar: open-source Rust LLM gateway for load balancing across multiple LLM endpoints (OpenAI, Anthropic, local).
DSA Trainer tool for LeetCode practice uses hint ladders instead of solutions, built as alternative to ChatGPT for learning data structures.
Microsoft open sources pg_durable for fault-tolerant, long-running SQL functions in PostgreSQL using durable execution patterns.
Open source terminal AI coding agent using Tree-sitter knowledge graphs for codebase context and multi-provider model support.