Incoherence in Goal-Conditioned Autoregressive Models
Mathematical analysis of incoherence in goal-conditioned autoregressive models, studying policy improvement through fine-tuning with online RL.
Mathematical analysis of incoherence in goal-conditioned autoregressive models, studying policy improvement through fine-tuning with online RL.
Theoretical analysis of diffusion models on discrete state spaces, establishing convergence guarantees for masked and random walk dynamics.
Tomographic Quantile Forests (TQF) for nonparametric uncertainty quantification in multivariate regression tasks.
Meta-probabilistic modeling framework for discovering latent structure across collections of related datasets using probabilistic graphical models.
Research on learnable Gray-Wyner networks for disentangling common and task-specific information in computer vision.
SAU method for machine unlearning in sparse LLMs via gradient masking and importance redistribution for privacy.
Research showing activation steering vectors in LLMs are fundamentally non-identifiable with large equivalence classes.
FIRE method for reinitialization in continual learning that balances stability and plasticity in neural networks.
Research on Natural Hypergradient Descent for bilevel optimization using Fisher information matrix as Hessian surrogate.
Evaluation of scaling laws for Chemical Language Models on downstream molecular property prediction tasks.
Skill routing system for LLM agents that identifies relevant skills from large ecosystems before planning or execution.
Framework using LLMs to automatically design reward programs for cooperative multi-agent RL systems with sparse task feedback.
DreamerAD uses latent world models for efficient RL in autonomous driving, compressing diffusion sampling 80x with visual interpretability.
ERL framework enabling LLM agents to self-improve through experiential learning from past interactions and reflective adaptation.
Neuro-symbolic method for process anomaly detection combining neural networks with domain knowledge from process mining.
Hierarchical indexing system for efficient fine-grained sparse attention in transformers, removing bottleneck from key selection.
Uses 2-datapoint reduced density matrix from quantum chemistry to predict and understand phase transitions during neural network training.
Continual learning framework using hierarchical exploration-exploitation to acquire knowledge from task streams without catastrophic forgetting.
Combines MCMC correction with score-based diffusion models using Metropolis-Hastings steps for improved sampling in model composition.
Method for estimating intrinsic dimensionality of datasets accounting for scale-dependent effects and measurement noise in unsupervised learning.
LLM-based approach for unsupervised code correctness evaluation that separates code comprehension from auditing to improve accuracy without reference implementations.
Project management framework using GenAI agents to optimize team composition by matching personality roles.
Addresses negative transfer in fine-tuning by selectively forgetting unhelpful pre-trained knowledge in language models.
Variance-based pruning method for compressing trained networks including Vision Transformers with minimal retraining.
NES framework for low-latency code edit suggestions without explicit instructions, using learned editing trajectories.
Open source CayleyPy library for efficient Cayley and Schreier graph computations, with 200+ new conjectures in group theory.
Retrieval-of-Thought (RoT) system reuses prior reasoning steps organized in thought graphs to improve LLM inference efficiency.
Evaluates self-replication risks in LLM agents through realistic testing of autonomous agent behaviors and safety concerns.
Proposes flow matching method for Bayesian posterior inference without likelihood evaluation, using block-triangular velocity fields.
RAG system for exhaustive multi-document question answering that checks all relevant documents without clear stopping conditions.
Multi-agent reasoning framework using AI agents for interpreting gene clusters in antimicrobial resistance transcriptomic data.
Framework using conformal prediction to assess correctness of LLM outputs and construct confidence sets for generative model responses.
Data-free quantization techniques for CLIP vision-language models enabling model compression without real data access for privacy-sensitive scenarios.
Study showing structured prompts significantly improve language model evaluation accuracy compared to single static prompt configurations in benchmarking.
LLM-based framework bridging cross-domain data sources for stablecoin transparency in circulation, reserves, and disclosure records.
RoboNeuron middleware layer connecting Vision-Language-Action models and LLM agents to robot middleware, standardizing tool API integration for embodied AI.
EvalBlocks modular framework for efficient evaluation of foundation models in medical imaging, reducing manual experiment tracking workflows.
Survey of meta-learning and meta-reinforcement learning methods enabling rapid task adaptation with minimal data, tracing DeepMind's adaptive agent research.
OPERA data pruning framework for efficient dense retriever adaptation, balancing quality-coverage tradeoff in domain-specific finetuning.
AI agents autonomously perform high energy physics analysis pipeline stages including event selection, background estimation, and statistical inference using LLMs.
LLM Router uses internal prefill activations for query-specific model selection, outperforming semantic routing by capturing model-specific failures.
CarbonEdge framework for carbon-aware deep learning inference on edge devices, optimizing for environmental impact alongside latency and throughput.
Agent2 open-source runtime for production AI agents with schema-to-API capabilities, auth, and provider routing.
Database optimizers become critical infrastructure when AI agents autonomously generate SQL queries.
3D semantic atlas of 188 constitutions using embeddings and UMAP for conceptual law search.
Hybro interoperability layer enables local and remote AI agents to coordinate in shared networks.
Trinity-Large-Thinking 398B sparse MoE model with chain-of-thought reasoning and agentic RL.
Trinity Large Thinking open-source reasoning model from Arcee AI optimized for agentic tasks.
Tutorial using Model Context Protocol server to turn NAS into self-hosted AI assistant.
Analysis of AI adoption risks: organizations mistaking temporary model limitations for safety assurance.