Quantitative convergence of trained single layer neural networks to Gaussian processes
Theoretical analysis of quantitative convergence of shallow neural networks trained via gradient descent to Gaussian processes.
Theoretical analysis of quantitative convergence of shallow neural networks trained via gradient descent to Gaussian processes.
Research on NVFP4 quantization approach for efficient LLM pretraining, reducing compute and energy requirements for frontier models.
VidGuard-R1 uses reasoning MLLMs and reinforcement learning to detect AI-generated videos with human-interpretable explanations.
Research on self-supervised novel view synthesis identifies transferability as key criterion for true NVS capability across video sequences.
Safety filtering for reinforcement learning using Control Barrier Functions to enforce dynamic safety constraints during training.
Multimodal foundation model for accelerating numerical simulation of stochastic differential equations via neural network-based error correction.
CoRPO: Adds correctness bias to GRPO reinforcement learning for improved reasoning and generalization in LLMs.
ObAct: Imitation learning framework for active vision in dual-arm robots using 3D Gaussian Splatting.
GRAND: Multi-agent path finding system using reinforcement learning for robot fleet task scheduling and dispatch.
ReFusion: Masked diffusion LLM combining parallel inference with autoregressive KV caching for faster generation.
AMPEND-LS: Agentic multi-persona LLM framework for multimodal fake news detection with evidence grounding.
Parallel Token Prediction: Framework for generating multiple LLM tokens in single forward pass via deterministic functions.
EmboTeam: LLM-based multi-robot task planning framework using behavior trees and PDDL for embodied AI.
AI agent system for continuous long-horizon egocentric video understanding from wearable devices.
Theoretical convergence analysis of Muon optimizer for nonconvex optimization problems.
LatentChem: LLM framework for chemical reasoning using latent representations instead of chain-of-thought text.
Revisits Laplace mechanism for DP-SGD in high dimensions using majorization theory for private training of large language models.
Multi-agent system translating jailbreak papers into executable modules for unified benchmarking of LLM robustness with reproducible evaluation.
Studies whether interpreter state persistence should be part of LLM agent training. Shows runtime persistence affects tool-augmented agent behavior.
System infrastructure for LLM-driven agentic ML pipeline search where agents autonomously generate, validate, and optimize ML pipelines over Python libraries.
Multi-model ensemble using LoRA fine-tuning for code comment classification across Java, Python, Pharo. Combines four transformer encoders via PEFT.
CI/CD quality gate tool detecting failures in AI-generated code including hallucinated packages and logic gaps.
Shell helpers piping git diffs to Claude API for automated code review and criteria generation.
SaaS tool using AI to prioritize feature requests from multiple sources by understanding user context.
Zalor platform for automated testing and scenario generation of AI agents before production deployment.
Open-source security scanner detecting vulnerabilities in AI coding assistant configurations (Cursor, Copilot, Cline).
BiomeSyn ecosystem simulator for testing long-horizon multi-agent AI behaviors with memory and cooperation.
Continuation of LLM-based reverse engineering: converting decompiled binaries to modern programming languages.
Using LLMs to automate binary decompilation and reverse engineering of compiled programs.
Bruce Perens argues AI will undermine copyleft licensing models citing chardet library license change.
Platform for building AI agents and autonomous workflows with integrations to 1000+ applications.
Graduate student releases MIT-licensed 3D C++ OpenGL engine built with agentic AI coding assistants to test their capabilities on complex systemic tasks.
Discussion thread on multi-agent AI system architectures and workflows, including 13-agent PAI Family example.
Polyscope: IDE designed for AI agent-first development. Limited details provided.
MCPSec scans Model Context Protocol configs for OWASP MCP Top 10 security risks. Developer tool for securing AI agent infrastructure.
Fractals: recursive task orchestrator for agent swarms using git worktrees and batch execution. Open-source AI agent framework.
SlideScholar converts research papers to conference slides via Claude API. LLM application with open-source stack (Next.js, FastAPI).
OpenAI Symphony: autonomous agents orchestrate project work from Linear board, execute tasks with CI/PR proof-of-work. AI agent framework.
CLI tool enabling Claude Code sessions to transfer between machines with local file access. Developer tool for AI coding.
AI agent running actual business as CEO with open-source codebase, public decision logging, and goal to reach $80k/month revenue.
CLI tool for GPU provisioning across 19 cloud providers with automatic vLLM optimization and Kubernetes deployment. Developer tool for LLM ops.
Luma's Uni-1 unified multimodal model for generation and understanding across image, video, audio, and text with agentic capabilities.
Video analyzing 20M GitHub PRs with Jellyfish to extract insights and benchmarks for AI development.
Platform renting idle browser instances to AI agents for web automation tasks, bypassing bot detection and CAPTCHAs.
Codex Fast Mode feature enabling 1.5x speed increase on GPT-5.4 at 2x credit cost. Developer tool documentation.
Anthropic research paper measuring and analyzing labor market impacts of AI with new methodology.
Open-source permissions and approvals framework for AI agents with SDK for enforcing boundaries, tracking actions, and user control.
Standalone verification tool for code changes and AI agent behavior with proactive issue detection after agent modifications.
Local-first knowledge graph for developers that watches project files, extracts entities using LLMs, and enables natural language querying.
AI tool that browses applications, generates human-readable test specs, then writes maintainable Playwright E2E tests.