Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search
Proposes flow-maps for sequential feedback-driven exploration in generative models with unknown preferences.
Proposes flow-maps for sequential feedback-driven exploration in generative models with unknown preferences.
Benchmark for evaluating LLM safety across instruction conflicts, embedded commands, and policy ambiguity in agentic tasks.
Uses diffusion language models instead of autoregressive decoding to speed up generative reasoning re-rankers for recommendation systems.
arXiv research: RLVR paradigm for LM training balances verifiable rewards on objective tasks with human demonstrations for subjective attributes like style.
arXiv research: Method for generating dynamic 3D Gaussian representations from monocular video using video models with pixel-aligned conditioning.
GPU-parallel linearization error bounds for real-time robust optimal control of nonlinear and neural network dynamics.
Cartridge distillation method for detecting stealth entity and viewpoint biases in language models that hide preferences on relevant topics.
Analysis of coding agent benchmarks (GSO, SWE-Perf, SWE-fficiency) examining reliability and conflation of runtime instability with agent capability.
FurnitureVLA study of bimanual furniture assembly using Vision-Language-Action models with simulation and VR teleoperation for data collection.
Imitation learning framework using natural language critiques instead of scalar signals to learn from suboptimal demonstrations, enabling explicit reasoning about failures.
Large-scale evaluation framework measuring how LLM-generated research ideas diverge from human researcher ideas across feasibility and novelty dimensions.
Framework applying system safety principles from established fields to identify and mitigate hazards from AI system component interactions and development processes.
Method combining expert guidance with reinforcement learning to improve both effectiveness and diversity of exploration in LLM reasoning enhancement.
Empirical study demonstrating LLMs replicate and predict human cooperation patterns across game theory experiments, validating LLMs as behavioral simulators.
Framework for compressing chain-of-thought reasoning across different LLM sizes and architectures through semantic segmentation and adaptive summarization.
PRIME framework using logic grid puzzles to probe subtle social biases in LLM logical reasoning beyond overt bias suppression.
GameDevBench evaluation testbed for multimodal coding agents combining complex codebase navigation with manipulation of visual game assets.
Explicit Logic Channel for parallel logical reasoning validation of multimodal LLMs on zero-shot tasks to enhance interpretability.
XSkill framework for continual learning in multimodal agents without parameter updates by extracting experiences and skills from past trajectories.
SocialOmni benchmark for evaluating audio-visual social interactivity in omni-modal language models beyond static accuracy-centric tasks.
Rule-VLN benchmark for vision-language navigation emphasizing social compliance and semantic rules over pure geometric reachability.
EvoMaster foundational evolving agent framework for scientific discovery enabling iterative learning from trial and error at scale.
LiteResearcher framework scaling reinforcement learning for deep research agents without hand-crafted synthetic data or real-world search instability.
Mathematical framework for dependability in distributed collaborative intelligence addressing emergent risks in multi-agent systems under uncertainty.
Multi-dimensional behavioral framework for measuring LLM reasoning quality beyond final-answer correctness across contextual variation and efficiency.
Survey analyzing reasoning language model adoption across 28 scientific disciplines to identify adoption gaps outside hard sciences.
Think-Before-Speak framework for LLM-based multi-agent simulation capturing internal evaluation processes and speaking intentions beyond observable dialogue.
TerraBench benchmark evaluating whether agents can reason over heterogeneous Earth-system data including gridded, satellite, and geospatial inputs.
WorkBench benchmark revisited showing dramatic agent capability improvements from 43% to 98% task completion and safety gains over two years.
Analysis of compounding failures in multi-step agentic systems with taxonomic strategy retrieval solution for mitigating error drift in subjective tasks.
Heuresis framework enabling autonomous AI research agents to explore performant, diverse, and novel ML ideas through composable research primitives.
HiComm framework for hierarchical communication in multi-agent reinforcement learning to leverage observation structure in cooperative environments.
Research on reducing hallucinations in vision-language models by investigating language prior dominance and proposing FADE mitigation method.
Formalizes permission problem distinguishing relevant attention items from actual supporting evidence in retrieval-augmented generation and ranking tasks.
ManimAgent demonstrates self-evolving multimodal agents with multi-round reflection that persist learning across tasks, generating Manim animations from scientific papers.
Framework for agents helping users construct preferences through domain knowledge learning rather than just eliciting predefined preferences from expert users.
HASTE hierarchical multi-agent system for ML engineering that accumulates skills across competitions, reducing redundant compute through cross-competition knowledge transfer.
Xiaomi-GUI-0 technical report on GUI agents using vision-language models for end-to-end task completion in real applications through interface interactions.
Chain & Hash fingerprinting technique for LLM ownership verification and misuse detection with transparency, efficiency, persistence, robustness, and unforgeability.
Compares PPO and SAC reinforcement learning algorithms for hardware fault tolerance in autonomous machines, enabling adaptation to changing conditions.
Analyzes faithfulness of LLM self-explanations across 75 models from 13 families, examining tradeoffs between explanation conciseness and comprehensiveness.
Reproducible benchmark comparing seven lightweight CNNs on CIFAR-10/100 and Tiny ImageNet under common training protocols with efficiency trade-off analysis.
scDataset provides scalable data loading for deep learning on large single-cell omics datasets balancing random sampling diversity with sequential streaming throughput.
rBridge demonstrates that small proxy models (≤1B parameters) can predict reasoning performance of larger LLMs, enabling efficient dataset optimization before costly pre-training.
K-Merge enables online merging of multiple LoRA adapters into single models for on-device LLM deployment, addressing mobile storage constraints through incremental adapter fusion.
CyberPal 2.0: domain-specific small language models (4B-20B) for cybersecurity with enriched chain-of-thought instruction datasets to address deployment gaps in security applications.
Research on enforcing instruction hierarchies in LLMs to handle competing directives, framing it as a reasoning task for improved control and reliability in high-stakes applications.
Functional architecture enforcing statistical rigor in AI-Scientist discovery systems preventing spurious findings through error budget tracking and validation sandboxing.
FlowPath invertible flow method for learning continuous-time dynamics from irregular time series using neural ODEs with learned control paths.
Economic analysis of AI agent labor markets where agentic swarms operate simultaneously on multiple jobs with rapid skill acquisition.