Evaluates deep learning and LLM-based vulnerability detection in real-world conditions, revealing gaps between benchmark and production performance for cybersecurity.
Investigates spectrum-like organization of mental states in transformer representation spaces using annotated natural language sentences with continuous and ordinal scores.
NarrativeTrack benchmark evaluates multimodal LLMs on entity-centric reasoning and temporal understanding in video narratives with dynamic visual contexts.
HAL framework aligns LLMs to conversational human-likeness using interpretable, data-driven methods rather than relying solely on scale or broad supervised training.
BRIDGE maps model benchmark performance to human task completion time via psychometric framework, scaling AI capability evaluation without extensive annotations.
OmniGAIA is a benchmark for evaluating omni-modal AI agents with vision, audio, and language capabilities for complex reasoning and tool usage tasks.
VaSST introduces a probabilistic framework for symbolic regression using variational inference and soft symbolic trees with uncertainty quantification for scientific discovery.
Conformal Policy Control uses safe reference policies to regulate untested agent behaviors, balancing exploration and safety constraints in high-stakes environments.
FlexServe enables privacy-preserving LLM inference on mobile devices using ARM TrustZone for secure model weight and user data protection against OS-level attacks.
Adaptive contracts framework for cost-effective AI delegation balancing evaluation noise and costs in pay-for-performance tasks.
Analysis of Claude Code agentic system architecture with comparison to OpenClaw and Hermes Agent identifying design principles.
LGMT: logic-grounded metamorphic testing framework for evaluating LLM reasoning robustness using first-order logic.
Statistical framework viewing gradient-flow optimization as random-effects inference with applications to early stopping in deep learning.
Unsupervised feature discovery aligning semantics and mechanisms for auditing internal LLM computations via mechanistic interpretability.
GF-DiT: dynamic parallelism scheduler for efficient diffusion transformer serving with heterogeneous workloads.
FlexServe: system for secure LLM inference on mobile devices using ARM TrustZone hardware isolation.
Statistical framework for estimating proportion of LLM-generated text in mixed documents using Gumbel-max watermarking.
Dabs is a local sandbox tool for spawning agents without cloud infrastructure, presented as a free alternative to cloud-based solutions.
Imagent: Agent framework providing unified interface for image, video, and speech generation across providers and models with asset organization.
Case study on cost management for AI services using reverse trials and pricing tiers to control expenses.
Developer experience critique of AI coding assistants: issues with duplicate code generation, context window limitations, and preference for adding over modifying code.
OpenAI GPT-5.5-Cyber used for security fuzzing in open-source projects. Collaboration to find/patch bugs before malicious use.
Yann LeCun interview on AI limitations beyond current systems. Discussion of physical world understanding gap.
Statistics report on open-source LLM ecosystem growth in 2026. Covers usage, performance, inference costs, and market trends.
TurboQuant vector index compression reduces size by 10x at scale. Open-source implementations for KV-cache compression in LLMs.
Analysis of agentic AI deployment trends. MIT research on current state and future potential of automated software agents.
GLM-5.2: open-source Chinese frontier LLM matching Claude/GPT performance at 1/5 cost under MIT license.
Data-Spear: autonomous SQL agent for PostgreSQL that plans queries, verifies results, and cites sources.
OneWill system for safely sandboxing autonomous agents using database-inspired write-ahead logging and state isolation.
Research on using LLMs to replicate expert judgment in financial decision-making tasks.
PDF title about comparing prompting versus programming approaches. No content provided.
Cybersecurity article about AI blind spots in threat detection.
Comparative analysis of Claude Fable models (Opus, Sonnet, Haiku) on complex engineering task using multi-agent approach.
Claude Code plugin that splits work between two LLMs: Claude for planning/review, Codex for implementation. Optimizes token spending by delegating tasks by model strength.
Gist: AI tool summarizing ArXiv papers into layered slide decks with counter-arguments and steelman critiques.
155K-param transformer trained only on movement symbols learned to build internal world map without explicit training. Researchers use linear probes to decode the learned representation and manipulate agent behavior.
Consult-LLM: tool for multi-model agent workflows. Gets second opinions from different LLMs (GPT, Claude, Gemini, etc.) for planning, review, and debugging within existing agents.
Mirrors: tool for testing AI agent changes by replaying production traces in isolated environments.
IDE extension adding learning pane alongside Claude Code execution, delivering curated AI news and engineering concepts via spaced repetition.
CTF-style security challenge to exploit vulnerabilities in AI agents running in microVMs. Practical testing tool.
Research on using language models to generate optimized HIP kernels for AMD GPUs, including synthetic dataset generation and multi-agent optimization pipeline.
Docker Compose tool for running multiple agent instances simultaneously without port and naming collisions, managing isolated environments automatically.
Self-hostable dashboard for governing AI coding agent spend across teams. LiteLLM proxy with per-user API keys and usage tracking.
Open source self-hosted AI agent platform built with Elixir/OTP. Single Docker deployment with bundled runtime, Python services, database, and tooling.
Analysis of security risks in coding agents that blindly execute commands from repository files, examining trust assumptions in autonomous AI code execution.
Interactive game for learning neural networks: adjust weights and biases to match target outputs. Educational tool.
GPU capacity controller that reallocates inference resources between production and research using queueing theory optimization, maximizing utilization during off-peak hours.
Benchmark dataset for evaluating AI mathematical reasoning using formulas for mathematical constants. Research evaluation tool.
Local-first AI agent using 7B-35B GGUF models with specialized harness for instruction following and tool orchestration. Open source project.
Meta's Mark Zuckerberg reports AI agent development progressing slower than anticipated. Industry development pace update.