Detection method for training data used in RLVR (reinforcement learning with verifiable rewards) by analyzing structural convergence of reasoning to identify benchmark contamination.
Lightweight entropy-guided framework predicts LLM output length using internal hidden states to reduce computational waste from excessive padding in batched inference.
REVIS framework corrects object hallucination in vision-language models by extracting suppressed visual information through latent space steering without training.
Prototype Transformer (ProtoT) introduces interpretable-by-design LM architecture using prototypes to make reasoning explicit and reduce opacity, hallucination, and deception risks.
Talk2DM system enables natural language querying of dynamic maps for vehicle-road-cloud integrated autonomous driving using LLMs for traffic scene representation.
Framework for AI agents to dynamically decompose complex tasks into sub-components and delegate to other agents/humans while adapting to environmental changes and handling failures robustly.
Hierarchical Sparse Autoencoder (HSAE) captures hierarchical structure in LLM features by building on sparse autoencoders, addressing feature splitting and monosemantic feature extraction.
Selective Abstraction framework enables LLMs to partially abstain on uncertain information in long-form generation rather than all-or-nothing, reducing factual errors while preserving useful content.
LLM-based approach for quantitative finance using logic-oriented paradigm to model financial market movements beyond asset-centric or market-centric methods.
Gaia2 benchmark for evaluating LLM agents in realistic asynchronous environments with temporal constraints, dynamic events, and multi-agent collaboration.
CSEval framework evaluating clinical semantic accuracy in medical text-to-image generation for anatomical and pathology correctness.
InjectRBP method steering LLM reasoning behavior through systematic pattern injection to enhance reasoning performance.
LawThinker autonomous legal agent with Explore-Verify-Memorize strategy for procedurally compliant reasoning in dynamic judicial environments.
Differentiable modal logic framework for debugging multi-agent AI systems by reasoning about knowledge, belief, causality, and obligation.
Pensieve Paradigm enabling stateful language models to autonomously manage and retrieve their own context using database operations.
Method for training large reasoning models with adaptive reflection and length penalty to reduce token consumption while maintaining accuracy.
Hadamard Linear Attention mechanism reducing computational cost of standard quadratic attention in transformers using kernel functions.
Quantitative analysis of gender and skin-tone bias in Gemini Flash 2.5 and GPT Image 1.5 using neutral prompts with rigorous measurement pipeline.
Value Alignment Tax framework measuring how alignment interventions propagate unintended effects across interconnected values in LLMs.
STAR framework combining statistical methods and agentic reasoning for predicting large model performance from limited observations with explanations.
Lossless data compression method using discrete latent transformers and reinforcement learning to exploit structure in complex data formats.
Evaluation framework testing whether GPT-4o possesses theory of mind through causal models of mental states, finding core features lacking.
Sci-CoE framework for co-evolving LLMs on scientific reasoning tasks using geometric consensus with sparse supervision for improved reliability.
Logical graphical model (QBBN) extended with negation and natural language parser for information retrieval using probabilistic reasoning.
Pedagogically-inspired framework for knowledge distillation from LLMs to smaller models using synthetic data with systematic learning process awareness.
Analysis of SAM3 text encoder for vision-language segmentation, examining over-provisioning in capacity for short, structured prompts.
Study of speech recognition model failures on high-stakes utterances (U.S. street names) revealing systematic transcription errors across 15 commercial models.
Physics-guided LLM agent for symbolic equation discovery modeling multi-step scientific reasoning process instead of direct data-to-equation prediction.
CM2 framework using checklist-based reward signals for multi-turn, multi-step agentic tool use via reinforcement learning without requiring executable environments.
CATTS technique for dynamically allocating compute in multi-step web agents using test-time scaling to improve agentic task performance and reliability.
Reinforcement fine-tuning approach for medical vision-language models integrating perception and reasoning augmentation for cross-modal tasks.
HybridRAG framework for LLM chatbots combining retrieval-augmented generation with pre-generated Q&A over unstructured documents for practical deployment.
LLM-based automated optimization modeling using error-driven post-training analysis to improve decision-making assistance with high-quality synthetic training data.
RECOM benchmark dataset of 15,000 Reddit questions evaluates how open-source LLMs (Llama3.1, Mistral) handle temporally recent open-domain QA tasks.
Study on parameter-efficient fine-tuning methods and their effects on LLM hallucination detection.
FalseCite dataset benchmarking LLM hallucination via internal state clustering across multiple models.
NLP research on text classification for SDG documents using combinatorial fusion and generative AI.
Analysis of transformer representations showing distinct functional roles for direction and magnitude in hidden states.
Bayesian optimization framework for efficient LoRA hyperparameter tuning in LLM fine-tuning using language-aided search.
Quantifies tokenization efficiency disparities across writing systems in multilingual LLMs, measuring inference latency impacts.
Evaluates few-shot temporal reasoning capabilities of LLMs for predicting human activities in smart home environments.
Fine-tuning and probing LLMs for Alzheimer's disease detection, analyzing task-relevant information encoding and data synthesis.
Survey of prompt engineering techniques for LLM applications, examining structured frameworks for NLP task performance.
Vector quantization compression method for Mixture of Experts LLMs using KLT-guided SVD for ultra-low-bit deployment.
Research on optimizer design for LLM training, analyzing spectral anisotropy in gradient signals to improve learning of contextual information.
Multi-agent reinforcement learning approach for optimizing chiplet placement in 2.5D circuits, handling thermal and wirelength constraints.
Time-TK: time series forecasting combining Transformers and Kolmogorov-Arnold Networks with multi-offset temporal interactions.
MELINOE: fine-tuning approach enabling memory-efficient inference for Mixture-of-Experts models through selective expert loading.
DDL2PropBank: benchmark for evaluating multi-agent framework developer experience through database schema to semantic role mapping task.
Defense method against downstream-agnostic adversarial examples in pre-trained SSL encoders without task-specific fine-tuning.