Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
Practical adversarial attacks on stochastic bandits via fake data injection with bounded perturbations, more realistic than prior per-round manipulation assumptions.
Practical adversarial attacks on stochastic bandits via fake data injection with bounded perturbations, more realistic than prior per-round manipulation assumptions.
Provably safe RL using analytic gradients for autonomous robots, integrating safety safeguards during training to reduce sim-to-real deployment gaps.
Discourse-aware hierarchical retrieval for long document QA using rhetorical structure theory, improving over flat chunking approaches for LLM comprehension.
Survey of federated foundation models for privacy-preserving recommendation systems, addressing integration of FMs in decentralized collaborative settings.
Attribution-guided pruning discovers circuits in small LLMs for mechanistic interpretability, enabling diagnosis and targeted correction of undesirable behaviors.
Multi-objective instruction-aware RL for procedural content generation leveraging natural language control, addressing limitations in handling complex textual instructions.
Improves Equilibrium Propagation learning framework with feedback regulation and residual connections for brain-inspired computing, addressing instability and computational costs.
Study of multi-step reasoning in LLMs using cellular automata framework, examining how recurrence, memory, and test-time compute scaling improve reasoning depth beyond memorization.
RFA reformulates self-attention as robust state estimation using linear SDEs, matching computational complexity of standard attention while improving theoretical foundations.
Frictional Q-Learning reduces extrapolation errors in off-policy RL by treating replay buffer as low-dimensional manifold, using friction analogy for action selection.
KLCF framework addresses hallucination in LLM long-form generation by incorporating knowledge-level consistency into RLHF, aligning model outputs with its own knowledge boundaries.
Unsupervised framework using optimal transport for learning procedures from instructional videos by handling noise and execution variability.
Framework for automatically searching the internet to construct challenging benchmarks at scale without human curation for model evaluation.
Proof-of-concept combining role-playing games with LLM analysis to elicit moral profiles of users in software requirements engineering.
Theoretical analysis of Reinforcement Learning with Verifiable Rewards (RLVR) for post-training LLMs using binary feedback, introducing Gradient Gap metric.
Studies inference-time scaling using pause tokens to improve foundation model expressivity while maintaining parallelizability for faster reasoning.
Comprehensive review of Kolmogorov-Arnold Networks as structured alternatives to MLPs, covering theory, relationships to classical methods, and applications.
Proposes AsyncVLA, a Vision-Language-Action model using asynchronous flow matching for improved long-horizon robotic task execution with self-correction capabilities.
Stellar VLA: continual learning framework for vision-language-action models that evolves skill knowledge without parameter expansion.
RobustSora: benchmark for detecting AI-generated videos, controls for watermarks to isolate genuine generation artifacts.
SoccerMaster: vision foundation model for soccer understanding tasks from detection to semantic reasoning.
MediEval: benchmark linking EHRs to medical knowledge base for evaluating LLM reasoning and reliability in medical applications.
Deep Bayesian RL framework with learnable basis functions for improved generalization in Meta-RL tasks.
Research mapping human anti-collusion mechanisms to multi-agent AI systems. Studies how autonomous agents develop collusive strategies.
IGBO framework trains interpretable ML models balancing accuracy and explainability using feature importance hierarchies and gradients.
CSMCIR method for composed image retrieval combining text and images using CoT-enhanced alignment and memory bank.
HERMES: training-free architecture for efficient streaming video understanding in MLLMs using hierarchical KV cache memory.
Study showing LLMs are state-blind: ignore contextual/situational factors while capturing trait-based personas. Introduces Chameleon dataset.
Research on ensemble methods for causal discovery algorithms with expert guidance. Machine learning methodology for practical applications.
Cap-and-trade regulatory framework proposal to incentivize AI efficiency and accessibility, addressing resource equity and sustainability.
LLM-AutoDP framework using LLM agents to automatically process and clean domain-specific data for model fine-tuning.
Gossip-based decentralized learning algorithms for edge devices with communication efficiency, robustness to corruption, and low memory.
Theoretical analysis of optimal Attention/FFN ratios in disaggregated LLM serving architectures for efficient resource provisioning.
Meta-evaluation framework enabling self-evolving LLMs for non-verifiable tasks through LLM-as-Judge with quality-aware training.
FIT to Forget method for robust continual unlearning in LLMs handling sequential privacy and copyright deletion requests.
Leviathan transformer architecture decoupling input embeddings and output projections via learned embedding vectorization.
Neural solver for vehicle routing problems using lifelong learning with continually drifting task patterns and limited training.
DialectLLM framework for generating multi-dialectal conversational data beyond Standard American English, addressing LLM dialect representation.
Analysis of forgetting illusion in concept erasure for diffusion models via latent variable optimization attacks.
Statistical membership inference method for reliably auditing machine unlearning in models, addressing right to be forgotten.
Workflow for systematic evaluation of audio description quality using human raters and vision-language models at scale.
Novel training method for time-series forecasting models incorporating autoregressive rollout and error-growth heuristics from LLM training.
Theoretical analysis of transformer capabilities on the PARITY task, examining fundamental computational limits of neural architectures.
Probabilistic framework formalizing code selection and generation paradigms for test-driven development with AI assistants, analyzing environment-interaction strategies.
Flow matching approach for diffusion-based robotic policies reducing inference latency through informed noise sampling.
Adaptive curriculum enhanced group relative policy optimization for autonomous ML engineering agents using RL.
Attribution evaluation framework for multimodal LLMs assessing grounding across heterogeneous sources and modalities.
Parallel reasoning approach for visual comprehension in LLMs shifting from depth to parallelism to improve exploration.
Framework integrating deep generative models with quantum annealing for molecular design beyond training data.
Transformer architecture rethinking temporal and channel dependencies in medical time series like EEG and ECG.