Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
Revisitable memory mechanism enabling long-context LLM agents to reference and reason over dispersed evidence across millions of tokens.
Revisitable memory mechanism enabling long-context LLM agents to reference and reason over dispersed evidence across millions of tokens.
Epsilon-scheduling technique for robust fine-tuning from non-robust pretrained models addressing adversarial robustness transfer.
Method conducting multiple controlled pretraining experiments simultaneously in single training run to reduce computational costs.
Analysis showing Group Relative REINFORCE exhibits off-policy properties, demystifying GRPO and related LLM RL algorithms.
Joint optimization of prompts and parameters for LLMs exploring synergistic effects between prompt engineering and fine-tuning.
SimuHome benchmark with 600 episodes for evaluating LLM-based smart home agents on temporal and environment-aware task execution.
Uni-X architecture mitigating gradient conflicts between vision and text modalities in unified multimodal models.
Vid-LLM video-based 3D multimodal LLM combining 3D reconstruction and reasoning capabilities for 3D scene understanding.
Study of loss curve collapse phenomenon during LLM family training under practical scaling conditions for predictable scaling laws.
EasySteer unified framework for efficient and extensible LLM steering through hidden state manipulation at inference time.
SpinBench diagnostic benchmark evaluating spatial reasoning and perspective-taking capabilities in vision language models.
Method to improve LLM confidence calibration using self-generated distractors to reduce overconfidence and improve safety.
Knowledge distillation method for efficient LLM inference using concrete score matching to preserve logit information beyond softmax smoothing.
COMRES-VLM framework using vision language models for coordinated multi-robot exploration and object search in unknown indoor environments.
EditReward human-aligned reward model for scaling synthetic training data in open-source instruction-guided image editing models.
AdaBlock-dLLM improves diffusion-based LLM inference efficiency through adaptive block-size selection for semi-autoregressive decoding.
MENLO framework for evaluating native-like LLM response quality across 47 languages using human-annotated preference pairs and audience design principles.
C³B benchmark for evaluating cultural awareness in multimodal LLMs using comics with cross-lingual tasks and varying difficulty levels.
arXiv paper analyzing loss of plasticity in deep learning on non-stationary data using dynamical systems theory.
arXiv paper on stabilizing policy gradients for sample-efficient RL in LLM reasoning to improve training efficiency.
GEM is an open-source environment simulator for training agentic LLMs through experience-based learning, analogous to OpenAI Gym.
arXiv paper proposing RLP, reinforcement learning as a pretraining objective for reasoning models instead of post-training only.
arXiv paper on ExGRPO, reinforcement learning from verifiable rewards to improve LLM reasoning with experience reuse.
arXiv paper showing benchmark contamination detection in reasoning models is easily evaded, highlighting evaluation vulnerabilities.
arXiv paper on untargeted jailbreak attacks against LLMs that optimize adversarial suffixes without fixed target responses.
Hierarchical preference learning for long-horizon LLM agents addresses granularity mismatch between trajectory and step-level signals.
Analysis of implicit models' expressive power showing infinite-depth weight-tied networks can match larger explicit models.
RACE Attention achieves linear-time complexity for long-sequence training, enabling multi-million token contexts.
Analysis of cross-entropy scaling law breakdown at very large LLM scales and investigation of contributing factors.
TiTok transfers token-level knowledge to enable LoRA parameters to work across different model architectures.
SwiReasoning enables LLMs to switch between latent continuous and explicit discrete reasoning for token efficiency.
NANOMIND hardware-software co-design framework for efficient multimodal model inference on battery-powered mobile devices.
Training method for LLMs to generate diverse parallel reasoning paths using global forking tokens for improved problem-solving.
Information-theoretic framework detecting emergent higher-order coordination in multi-agent LLM systems.
MorphArtGrasp enables cross-embodiment dexterous hand grasp generation using morphology-aware eigengrasp framework.
Reference-Grounded Skill Discovery algorithm for scaling unsupervised skill learning to high-dimensional agent control.
ChainMPQ addresses relation hallucinations in vision-language models using multi-perspective question chains.
Relational Transformer architecture enables zero-shot transfer across relational databases with varying schemas and structures.
Value Flows framework extends distributional RL by modeling return distributions with continuous flow-based methods.
Test-time optimization framework improving video generation model performance on compositional scenarios without retraining.
DISCO method reduces computational cost of ML model evaluation by condensing datasets while preserving benchmark validity.
Study of adversarial attacks against LLM-based monitoring systems used in AI control protocols for autonomous agents.
Using LLMs to test JavaScript obfuscators for semantic correctness preservation rather than just deobfuscation resistance.
World models for evaluating and improving generalist robot manipulation policies with unfamiliar objects using simulation rather than costly real-world rollouts.
GAR combines generative adversarial networks with reinforcement learning for formal theorem proving in Lean, addressing limitations of fixed problem sets in training.
Characterizes reasoning capabilities of masked diffusion language models by connecting to computational complexity theory.
Systematic methodology for developing reliable fine-grained evaluators of LLM-generated natural language math proofs.
Data-driven real-to-sim framework generating diverse high-fidelity urban environments from city videos for training embodied AI agents.
PolySkill framework enabling LLM agents to learn generalizable reusable skills across different websites and tools during interaction.
Survey and interviews characterizing digital companionship as human-AI relationship using ChatGPT and Replika for task assistance and social connection.