Zero-Shot Coordination in Ad Hoc Teams with Generalized Policy Improvement and Difference Rewards
Zero-shot coordination method for ad hoc multi-agent teams leveraging all pretrained policies for transfer without retraining.
Zero-shot coordination method for ad hoc multi-agent teams leveraging all pretrained policies for transfer without retraining.
Multi-agent reasoning framework for systematic prompt optimization in LLMs with interpretable score-aware improvement guidance.
Automated algorithm design paradigm for auto-tuning optimizers across diverse irregular search spaces in high-performance computing.
Framework using geo-foundation models for flood hazard mapping from SAR satellite imagery in data-scarce regions.
Embodied world model using multi-view trajectory videos to improve consistency in action-to-movement translation for robotic prediction.
Inverse reinforcement learning approach using LLM guidance and language to disambiguate robot rewards from demonstrations.
Knowledge-based VQA system augmenting multimodal LLMs with external document retrieval and reasoning for domain-specific visual questions.
One-shot adaptation framework improving vision-language-action model generalization across novel camera viewpoints and visual perturbations.
Systematic framework for enterprise knowledge retrieval using LLM-generated metadata to enhance RAG system performance.
Vision-language model combining linear and sparse attention for efficient unlimited-input processing with high-frequency visual perception.
Alternative activation functions surpassing Dynamic Tanh for stable training in normalization-free transformer architectures.
Learning theoretic study on extracting features in superposition from complex ML models, advancing interpretability research.
Defense mechanism against backdoor attacks in instruction-tuned LLMs using defensive poisoning to merge and break adversarial triggers.
Architecture separating energy-based world models from language generation in LLMs to improve semantic understanding beyond fluent text production.
Agentic APR system that dynamically generates bug reproduction tests alongside AI-generated code fixes to increase developer confidence.
Generalist value model for policy evaluation in LLM training with actor-critic methods, reducing computational cost of critic networks.
Large-scale study of training vision-language models up to 344K context length with open-weight model recipes and data pipeline analysis.
Analysis of trade-off between generative and understanding capabilities in multimodal models, proposing R3 framework to address conflict.
Reinforcement learning-steered diffusion model for neural architecture search on directed acyclic graphs, improving NAS efficiency.
Study of how LLM-scaffolded programming affects novice skill acquisition, proposing metacognitive guardrails to prevent cognitive outsourcing.
Study comparing automatic similarity metrics versus LLM-as-a-judge evaluation for clinical dialogue, using domain-adapted Llama-2-7B with LoRA.
Randomized experiment on law students studying targeted training interventions to improve productive use of LLMs in legal analysis tasks.
Error enumeration as reward approach for reference-free RL post-training in virtual try-on tasks with multiple valid outputs.
Study on using finetuned lightweight LLMs for topic-conditional sentiment extraction from news to forecast aluminum commodity prices.
Position paper framing multi-agent memory as computer architecture problem, proposing three-layer memory hierarchy and identifying protocol gaps for collaborative agents.
AgentDrift reveals unsafe recommendation drift in tool-augmented LLM agents when tool outputs are corrupted, exposing gaps in ranking-metric-based evaluation.
Sample-efficient hypergradient estimation for decentralized bi-level reinforcement learning in strategic decision-making problems.
Analysis of AI query approximation using lightweight proxy models to reduce cost and latency of LLM-based SQL queries on structured and unstructured data.
InCoder-32B is a 32B code foundation model optimized for industrial scenarios requiring hardware semantics, specialized constructs, and strict resource constraints.
Research investigation into how LLMs internally compute verbal confidence scores and whether they are generated just-in-time or cached during answer generation.
Method for inducing sustained creativity and diversity in LLM outputs during exploratory search tasks through iterative refinement techniques.
FedRG addresses performance degradation in federated learning with noisy annotations using representation geometry for improved noise detection.
ContractSkill framework converts web agent skills into executable artifacts with explicit structure for detection and local repair, improving agent reliability.
LLMON is a markup language designed to convey structured and semantic information to LLMs beyond plain text, improving prompt clarity and LLM interface usability.
KARMA framework addresses knowledge-action gap when fine-tuning LLMs for personalized search tasks at scale, proposing regularization methods for industrial recommendation systems.
Activation watermarking technique for detecting adversarial attacks on LLMs that evade safety monitoring while eliciting unsafe outputs.
GNN layer with per-edge routing for heterophilous graphs, comparing cost-sensitive aggregation against uniform spectral approaches.
Vision-language model for sleep staging from polysomnography waveforms generating AASM-compliant clinical rationales with auditable reasoning.
Controlled evaluation of how LLM model choice, size, and prompting strategies affect political text annotation, challenging conventional wisdom.
GNN technique using cross-attention and cohesive subgraph embedding to address oversquashing problem in graph neural networks.
3-bit weight quantization method for LLMs using rotation-domain smoothing via Fast Walsh-Hadamard Transform, improving precision in extreme quantization.
Benchmark dataset for evaluating vision-language models on Japanese scene text understanding, addressing multilingual complexity challenges.
EvidenceNet framework uses LLM-assisted pipeline to extract structured, evidence-grounded biomedical findings from full-text literature into knowledge graphs.
Adaptive resolution framework for multimodal LLMs that reduces visual token overhead through input-side compression before encoding.
Framework for post-training compression of generative AI models with single-line implementation, addressing quantization and calibration challenges.
Derives time-varying momentum schedule for neural network training from physics principles, eliminating need for manual tuning.
Hybrid CPU-GPU framework combining differentiable optimization with ILP solving for combinatorial scheduling.
Multi-agent LLM framework for Bayesian optimization exploring exploration-exploitation trade-off through implicit reasoning.
LLM agents for GPU kernel optimization using domain-specific language and speed-of-light guidance to reduce design space.
Amortized analog circuit generation system combining graph VAE and flow-matching models with SPICE validation.