Structured Agent Distillation for Large Language Model
Framework for compressing large LLM-based ReAct agents into smaller student models while preserving reasoning and action consistency.
Framework for compressing large LLM-based ReAct agents into smaller student models while preserving reasoning and action consistency.
Framework for generative modeling with enforced physical constraints using split augmented Langevin sampling for scientific applications.
Hierarchical differential model for inferring system degradation from sensor data by disentangling slow and fast temporal dynamics.
Text-trained LLMs perform zero-shot extrapolation of PDE dynamics, revealing three-stage in-context learning mechanism for spatiotemporal forecasting.
Mathematical study of Busemann functions in Wasserstein space with applications to geometric machine learning and data slicing.
Development of conformal prediction method that ensures counterfactual fairness in prediction sets for fair decision-making under uncertainty.
Zeroth-order optimization approach for continual learning that improves memory efficiency and addresses plasticity-stability tradeoffs without gradient computation.
Research on unifying in-context learning and activation steering as instances of a broader framework using belief dynamics to control LLM behavior at inference time.
Introduces RAT+, structured dilated attention architecture that enables sparse inference while maintaining long-range connectivity and accuracy.
Proposes controllable exploration strategy for RLVR training of multi-modal LLMs to address entropy collapse and policy degradation.
Introduces FlashOptim, memory-efficient optimizers for mixed-precision neural network training reducing per-parameter memory requirements.
Shows preference labels in LLM-as-judge training can function as covert communication channels, challenging assumptions about semantic supervision.
Investigates tokenizer pretraining impact on physics foundation models for emulating complex multiphysics phenomena in data-limited settings.
Structure-aware set transformers with temporal and variable-type attention for asynchronous clinical time series in EHR data.
Analyzes how MDP design choices (state composition, rewards, dynamics) affect sim-to-real transfer in reinforcement learning for industrial control.
Instance unlearning method for diffusion models removing specific outputs without text prompts, addressing unpromptable undesired generations.
Proposes iterative selection of Gaussian mixture priors to prevent posterior collapse in variational autoencoders.
AI system analyzing police bodycam footage at scale to assess officer-public interactions and improve government accountability.
Generative Predictive Control method augments frozen diffusion policies with action-conditioned world models for improved robot control without retraining.
Applies multi-agent reinforcement learning to greenhouse gas offset credit markets for emissions control and carbon project trading simulation.
Data-driven survey identifying 14,648 papers on LLM limitations from 2022-2025 using automated classification and expert validation across 250,000 academic papers.
Novel algorithm for multi-agent reinforcement learning using uncertainty quantification and selective exploration to improve sample efficiency in joint action spaces.
Research paper on measuring whether LLMs comprehend user intent beyond surface-level text patterns, addressing training-inference gaps in language models.
Reinforcement fine-tuning approach for LLMs applied to point-of-interest recommendation with improved semantic indexing.
Open benchmark suite comparing paired encoder and decoder architectures for NLP tasks with controlled parameter counts.
Adapter parameters for efficient multi-task LLM inference on-device via task merging for compositional learning.
Agentic Design Review System orchestrates multiple AI agents to collaboratively analyze graphic designs with meta-agent coordination.
arXiv paper analyzes theoretical limitations of embedding-based retrieval for diverse tasks including reasoning and code generation.
DiDi-Instruct trains fast few-step language generation via distillation from discrete diffusion LLMs while maintaining quality.
CodeEvolve open-source framework combines LLMs with evolutionary algorithms to synthesize optimized algorithmic solutions guided by execution feedback.
Jr. AI Scientist is an autonomous research system that mimics novice researcher workflows, conducting autonomous exploration and including risk analysis.
Novel text-only adaptation method for LLM-based ASR systems using text denoising to preserve speech-text alignment without fine-tuning disruption.
Multi-agent reinforcement learning system for width-scaled information seeking, exploring complementary depth scaling in LLM deployment.
LatentChem: Latent reasoning interface for chemical LLMs, decoupling chemical reasoning from text tokens for improved efficiency.
Study of phonological vector arithmetic in self-supervised speech models across 96 languages, analyzing representation structure.
Evaluation of small language models for role classification in human-robot interaction with zero-shot and one-shot adaptation.
Retrieval system learning node-specific Riemannian metrics on citation graphs for geometry-aware semantic search.
Systematic evaluation of LLM-based AI agents in Byzantine consensus games, testing agreement behavior in adversarial settings.
Study of reasoning techniques in LLMs for political opinion modeling and alignment with individual preferences.
Analysis of chain-of-thought reasoning in LLMs, comparing activation probing and early stopping across DeepSeek-R1 and GPT-OSS models.
Caching optimization for concept learning in description logic knowledge bases using supervised learning.
Differentiable equilibrium blocks for multi-agent incentive design in game theory and economics applications.
Mixture-of-Experts architectures for machine learning interatomic potentials with analysis of routing strategies and sparse activation.
Research into theoretical mechanisms of LLM phenomena: semantic prompt comprehension, in-context learning, and chain-of-thought reasoning.
Dataset creation using Wikidata to detect sociocultural biases in LLMs, focusing on Latin American languages and cultures.
LightPanda is a headless browser built for AI agent workloads with optimized performance, security, and CDP support.
ClawJetty adds live status pages and progress tracking to AI agents with shareable run links and task completion visibility.
Prowl discovers and ranks AI services/APIs via natural language queries. Claude tests services across 8 performance dimensions.
Developer framework enabling AI agents to build analytics applications on ClickHouse. Provides specialized interfaces for coding agents.
Research from Google on training LLMs to reason using Bayesian inference for improved probabilistic world modeling in agentic systems.