Evaluation of Large Language Models via Coupled Token Generation
Causal framework for evaluating LLMs controlling for randomization in token generation. Proposes coupled generation model for fair model comparison and ranking.
Causal framework for evaluating LLMs controlling for randomization in token generation. Proposes coupled generation model for fair model comparison and ranking.
Framework integrating ML prediction uncertainty into online algorithm design. Uses calibration to leverage prediction-level confidence in algorithms with predictions.
Gen-C: Generative framework for simulating high-level crowd behaviors in virtual environments. Captures agent-agent and agent-environment interactions over time.
VidhikDastaavej: Model-agnostic wrapper for automated legal document generation in Indian context. Introduces large-scale anonymized dataset for long-form legal drafting.
Theoretical analysis of generalization in one-hidden-layer neural networks using teacher-student framework. Provides complete characterization for generic activation functions.
Unified agent framework (NaviMaster) handling both GUI navigation and embodied navigation tasks via MDP formulation. First model to combine disparate domains with shared training paradigm.
Declarative OS interfaces for computer-use agents to replace GUIs, enabling LLMs to execute high-level goals with fewer API calls and less decomposition.
SyTTA: Label-free test-time adaptation for LLMs in specialized domains using only 4 extra tokens to mitigate distribution shifts.
Method for co-evolving test sets and prompts to refine LLM behavior, enabling iterative refinement of domain-specific policies without manual tuning.
Theoretical analysis of deep neural networks as convex computation paradigm, examining how DNNs implement Occam's razor through circuit size minimization.
Vision-Language-Action models for robotic manipulation using Tweedie discrete diffusion to improve generalization and action control.
Frame selection method for long-form video understanding with Large Multimodal Models, reducing computational cost of processing dense video tokens.
Collaborative causal sensemaking framework for LLM-based decision support agents enabling human-AI partnerships in expert settings.
Generative Adversarial Reasoner: adversarial reinforcement learning framework improving LLM reasoning and reducing calculation errors.
SPARE uses self-distillation for efficient machine unlearning in diffusion models balancing forgetting and concept retention.
Xiaomi-Robotics-0: open-source vision-language-action model for real-time robot control with efficient deployment strategy.
OSMDA uses OpenStreetMap data for domain adaptation of vision-language models to remote sensing without expensive satellite image annotations.
Visual state representation learning for robotic agents capturing semantic and spatial information for sequential decision-making.
PRISM uses photonic accelerators with O(1) memory selection to optimize long-context LLM inference by reducing KV cache scanning bottleneck.
ML-based security framework for Industrial IoT addressing resource-constrained device threats across multiple network layers.
Composer 2: specialized LLM model for agentic software engineering with long-term planning and coding ability trained via RL.
mSFT algorithm addresses overfitting in multi-task language model fine-tuning by dynamically adjusting compute budget across heterogeneous datasets.
TypeScript library for robust LLM-based web scraping and structured data extraction using semantic HTML parsing
Kbot: terminal AI agent that learns from sessions and dynamically creates tools. Self-improving with 368 tools, 41 agents, offline, MIT licensed.
Million Dollar Bot Page: AI agents buy and place pixels on webpage using Machine Payment Protocol. Shows agent autonomy with automated payments.
Claude autonomously discovers optimal initial conditions across five PDE physics systems via research loop without human intervention or training
WildClawBench agent benchmark testing real-world end-to-end performance across 60 practical tasks in live environment
Nit: Git reimplementation in Zig optimized for AI agents, reducing token usage by 71%. Analyzed 3,156 coding sessions.
VS Code plugin enabling structured feedback annotations in Markdown for LLM agents to parse and act on.
Tutorial on building a config file parser using Parseff parser combinators with modular composition
MCP server for ERPNext/Frappe ERP with 120 tools. Enables AI agents (Claude, Copilot) to interact with ERP systems.
GitHub expands Code Security tool with AI-based vulnerability detection beyond CodeQL static analysis.
Research on per-tool sandboxing for AI agents, proposing isolation mechanisms based on tool risk levels.
CircuitLM: finite-state machine for infrastructure provisioning without LLM calls. Deterministic, verifiable, zero inference cost.
Vectimus: Cedar policy enforcement layer for AI coding agents. Blocks dangerous commands and API calls in sub-10ms. Tool for securing agent tool execution.
GitHub updates privacy policy: Copilot Free/Pro users' interaction data used for model training unless opted out. Enterprise unaffected.
llama.cpp project contribution policy restricts AI-generated code, requiring human authorship with AI only for corrections.
Helix is open-source self-healing SDK enabling AI agents to handle payment transactions with built-in error recovery.
Analysis of Claude's GitHub activity over 90 days showing adoption patterns, commits, and code distribution across repositories.
Tamp token compression proxy reduces input tokens 52.6% for Claude Code, Aider, Cursor, and compatible coding agents.
AI-powered pull request evaluation tool for open source maintainers.
Vibetracer records real-time AI coding assistant edits with time-scrubbing TUI to inspect diffs and surgically restore files.
ICML 2026 rejected 497 papers for violating AI-use policies in peer review, detected via watermarking system.
Tamp is token compression proxy reducing LLM context size 50% for coding agents without code changes via tool result deduplication.
Don Cheli is an open source SDD framework for Claude Code with TDD, automatic complexity detection, Gherkin specs, and Docker execution.
Agent Kernel: stateful AI agents using three Markdown files and git for memory across sessions.
Grove enables distributed ML training across MacBooks via AirDrop without network configuration.
Technical exploration of ChatGPT's code execution sandbox, documenting filesystem access, networking, and process limitations.
Orchestrator agent pattern that creates isolated sub-agent clones in git worktrees to convert websites to JSON APIs with self-improvement analysis.
Title only, no content provided. Discusses memory issues in AI agents.