Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
Analysis of multilingual reasoning gaps in reasoning language models, showing deficits stem from language understanding failures in low-resource languages.
Analysis of multilingual reasoning gaps in reasoning language models, showing deficits stem from language understanding failures in low-resource languages.
Method for interpreting LLM reasoning by resampling multiple chain-of-thought branches to measure causal influence and underlying computation.
LLM-guided decompilation framework using context to improve re-executability of decompiled binaries for security analysis.
SynthAgent: Framework for web agent adaptation using synthetic data generation with quality filtering to handle hallucinations and trajectory noise.
GroupRank: Efficient passage reranking paradigm using LLMs with groupwise ranking to balance efficiency and accuracy.
LiveCLKTBench: Benchmark pipeline for reliably measuring cross-lingual knowledge transfer in multilingual LLMs with time-sensitive queries.
Framework for process-centric evaluation of agentic software systems, analyzing execution trajectories and reasoning beyond outcome metrics.
Theoretical framework for sparse dictionary learning in neural networks, analyzing piecewise biconvexity and spurious minima in mechanistic interpretability.
WisPaper: AI agent system for academic paper discovery and organization, addressing semantic search and workflow fragmentation challenges.
VPR-AttLLM framework using LLM semantic reasoning to improve geo-localization of crowdsourced flood imagery.
Multimodal RAG system enhanced with knowledge graphs for audio-visual retrieval, extending LLM capabilities to multimodal domains.
Research on variance-aware tree policies for Monte Carlo Tree Search, improving upon UCB-based methods used in AlphaZero-style algorithms.
CricBench: benchmark for evaluating LLMs on multilingual cricket analytics and domain-specific Text-to-SQL tasks.
Research questions whether small proxy model training reliably guides data curation decisions for full-scale frontier AI model pretraining.
Disco-RAG improves retrieval-augmented generation by capturing discourse structure and synthesizing knowledge from dispersed evidence.
Enhanced-FQL(λ) reinforcement learning framework with fuzzy eligibility traces and interpretable fuzzy rules for continuous control.
Defensive poisoning technique merges triggers to remove backdoors in instruction-tuned LLMs vulnerable to data poisoning attacks.
HAERAE-Vision benchmark with 653 real-world underspecified visual questions reveals vision-language model limitations with informal queries.
GanitLLM: Bengali mathematical reasoning model with difficulty-aware curriculum-based GRPO training pipeline.
Coverage-enhanced latent actions framework for controlling multimodal conversational agents with reinforcement learning.
Game-theoretic analysis of how expanding AI agent capabilities affects strategic interaction in bargaining, negotiation, and persuasion.
EZ-MIA: Training-free membership inference attack against fine-tuned language models to audit privacy risks from data memorization.
Cross-modal domain adaptation approach transferring image dataset knowledge to LiDAR for synthetic training data generation.
Analyzes parallelism and generation order in Masked Diffusion Language Models across 8 models and 58 benchmarks.
Uses persona-based evaluation with LLMs to support inclusive cycling infrastructure design by simulating diverse user experiences.
Systematic analysis of demographic bias in LLM-generated targeted messaging across GPT-4o, Llama-3.3, and Mistral-Large models.
MERMAID: Multi-agent system for fact-checking using LLMs with memory-enhanced retrieval and iterative reasoning to assess veracity of claims.
Agent memory system beyond RAG addressing agent-specific needs: bounded coherent dialogue retrieval with decoupling and aggregation.
Unified framework explaining LLM steering methods (fine-tuning, LoRA, activation interventions) as dynamic weight updates from control signals.
El Agente Estructural multimodal agent for autonomous molecular geometry generation and manipulation using natural language and vision.
Fake-HR1 hybrid-reasoning model for synthetic image detection balancing chain-of-thought reasoning with computational efficiency.
Pyramid MoA hierarchical mixture-of-agents architecture with decision-theoretic router for cost-optimized anytime LLM inference.
Gome agent for machine learning engineering using gradient-based optimization instead of tree search, scaling LLM-based reasoning.
SteerEval benchmark for evaluating LLM controllability across language features, sentiment, and personality at multiple specification levels.
Analysis of prompt injection attacks as role confusion where models infer text source by content style rather than origin.
Survey of resource consumption threats in LLMs including excessive generation attacks, resource efficiency requirements, and mitigation strategies.
Study of LLM alignment evaluation focusing on routing from concept detection to behavioral policy, using Chinese language models as case study.
Curriculum learning framework using cross-entropy games to automatically build general capabilities and discover skills in language models.
Introduces Step-Level Reasoning Capacity metric and LC-CoSR training method to measure and reduce reasoning rigidity in chain-of-thought reasoning.
Training-free spatial-temporal token compression method for video MLLMs achieving high-ratio visual token reduction via forest modeling.
Memory-sparse attention mechanism enabling LLMs to scale context to 100M tokens through efficient memory modeling instead of full attention.
Multimodal deception detection using schema-driven approach with audiovisual analysis across multicultural datasets for forensics applications.
LLM-enabled threat hunting framework for SOC analysts integrating Splunk with policy-guided decision making for APT detection.
Benchmark framework for evaluating multimodal LLMs as perceptual backbones for autonomous agents in 3D environments with decision-dense scenarios.
Reinforcement learning framework enabling multimodal LLMs to autonomously crop and focus on regions of interest for improved perception in complex visual scenes.
Coarse-to-fine reasoning framework using reinforcement learning for interpretable multimodal sentiment analysis with MLLMs and hint-guided training.
Bio-inspired self-evolving network architecture for autonomous agents using evolutionary approaches instead of static human-defined protocols.
Framework analyzing agent communication protocols for LLM systems across three layers: communication, syntactic, and semantic. Systematically organizes 18 representative protocols.
Evaluates whether LLMs can infer causal intervention effects from natural language descriptions using behavioral simulation on climate-psychology interventions.
LitPivot: tool supporting iterative research idea development through dynamic literature contextualization and AI-driven critique.