Effective Dataset Distillation for Spatio-Temporal Forecasting with Bi-dimensional Compression
Dataset distillation technique for spatio-temporal forecasting using bi-dimensional compression to reduce training data and model complexity.
Dataset distillation technique for spatio-temporal forecasting using bi-dimensional compression to reduce training data and model complexity.
Analysis of mean bias effects in FP4 quantization during LLM training, addressing numerical instability from anisotropic weight distributions.
Physics-informed neural networks framework for multi-task learning of diverse Navier-Stokes equations addressing negative transfer and shared principles.
End-to-end system for speaker-attributed speech recognition in multi-party conversations with temporal boundaries and speaker identity consistency.
Methods for aligning generative search systems with user preferences while ensuring robustness, safety, and handling noisy retrieval.
Multi-agent negotiation framework for aligning LLMs with conflicting values across multiple stakeholders using deliberative processes.
Study showing how exposure of generative AI capabilities through user interfaces undermines deepfake detection methods using benign prompts.
Multi-agent reinforcement learning system for coordinating UAV fleets in time-critical medical supply delivery under uncertain conditions.
SCORE recurrent architecture replacing layer stacking with contractive ODE-inspired updates for efficient deep neural network training.
Reinforcement learning method extending verifiable rewards to general reasoning domains using conditional expectation for free-form LLM answers.
Novel approach to detect and eliminate neural network backdoors using active paths with application to intrusion detection systems.
Addresses robustness of neural text-to-SQL models against database schema evolution to maintain performance on dynamic schemas.
Investigates structured linked data and knowledge graphs as memory layer for agent-orchestrated retrieval-augmented generation systems to improve accuracy.
Framework for estimating aleatoric and epistemic uncertainty in deep learning models using single model for high-stakes applications like medical diagnosis.
Proposes risk-aware evaluation framework for LLM red-teaming specific to financial services domain, capturing failure modes in regulated BFSI settings.
Evaluates whether speech-aware LLMs encode speaker identity and proposes a model-agnostic scoring protocol for speaker verification on API and open-weight models.
BALD-SAM applies disagreement-based active prompting to Segment Anything Model for iterative interactive segmentation refinement.
Analyzes reliability of cue-conflict benchmark for measuring visual bias in neural networks, finding instability in stylization-based methods.
Value-driven memory approach enables LLM-based kernel synthesis for NPU programming without expensive fine-tuning on data-scarce domains.
Generalist Value Model V0.5 serves as prior for sparse RL rollouts in reinforcement learning with verifiable rewards.
Bilingual extreme multi-label text classification dataset with GND taxonomy for library cataloging and agent-assisted indexing.
Parameter-efficient Diffusion Transformer generates synthetic DNA regulatory sequences with reduced memorization and faster convergence.
Dynamics-Predictive Sampling improves RL finetuning of reasoning models through active learning-based prompt selection strategies.
LookaheadKV optimizes KV cache management in LLMs for long-context tasks by predicting future token importance without generation.
Studies fine-tuning LLM backbones for text-to-speech systems, using LoRA for voice consistency and speaker-specific characteristics.
Historical Consensus method prevents posterior collapse in VAEs through iterative selection of Gaussian mixture priors.
Safe RLHF approach using stochastic dominance for distributional risk control beyond expected cost constraints.
Contact Coverage-Guided Exploration method for general-purpose dexterous manipulation using deep reinforcement learning.
GroundCount framework combines Vision Language Models with object detection to reduce hallucinations in counting tasks.
Overview of how AI can improve software engineering practices within Agile development frameworks and team workflows.
Examines methodological challenges and solutions for human uplift studies evaluating frontier AI systems using RCT methodology.
Study comparing how Vision Language Models recognize artistic style versus art historians' interpretations through interpretability research.
Automated system generating comedic sketch videos using population of LLM-based agents with different roles competing and improving iteratively.
Multi-agent systems of LLMs and neural networks solving problems through natural language dialogue, overcoming limitations of single models.
Extends intelligent tutoring system with personalized hint explanations based on student traits like Need for Cognition and Conscientiousness.
Generates interpretable control policies by representing them as Python programs, using LLMs with evolutionary algorithms and systematic evaluation.
Method using multiple pre-trained models with consistency-based abductive reasoning to handle distributional shifts in novel environments.
Research on limitations of RL for LLM reasoning. Proposes interleaved online fine-tuning to help models acquire new capabilities beyond base model knowledge.
Research paper introducing Yokai Learning Environment for zero-shot multi-agent coordination benchmarking, advancing beyond Hanabi Learning Environment.
STRIPS Transformer: Architecture enabling next-token prediction to yield world models supporting planning via symbolically aligned transformers for STRIPS action domains.
RADAR: Dynamic routing mechanism for reasoning LLMs optimizing performance-cost tradeoff by selecting appropriate model size and reasoning budget for deployment.
BiasBusters: Benchmark and methods for detecting and mitigating tool selection bias in LLM agents when choosing functionally equivalent providers from marketplaces.
CostNav: Economic navigation benchmark for physical AI agents evaluating real-world commercialization constraints through comprehensive cost-revenue analysis.
IndiMathBench: Human-verified benchmark for mathematical theorem proving with AI-powered human-assisted autoformalization pipeline to address high-quality training data scarcity.
Closed-loop molecular discovery system combining language models, property alignment, and strategic search for de novo drug ligand design and virtual screening.
Hierarchical curriculum learning approach for lifelong agents in Dark Souls III, decomposing combat control into five reusable skills via directed skill graphs.
MemOCR: Multimodal memory agent improving long-horizon agentic reasoning by allocating memory efficiently within limited context windows using layout-aware visual compression.
Multi-domain RL framework for LLMs using Reinforcement Learning with Verifiable Rewards, addressing collaboration of RLVR across different domains for expert-level performance.
Autonomous AI analysts built on LLMs that replicate many-analyst studies cheaply, quantifying how analytic decisions affect empirical conclusions on the same dataset.
Minimal agent for automated theorem proving with iterative proof refinement, library search, and context management for systematic architecture comparison across frontier models.