Improving LLM Reasoning with Homophily-aware Structural and Semantic Text-Attributed Graph Compression
Improves LLM reasoning on text-attributed graphs through homophily-aware structural compression to overcome context window limitations.
Improves LLM reasoning on text-attributed graphs through homophily-aware structural compression to overcome context window limitations.
Paper2Rebuttal multi-agent framework assists with transparent author response writing using verifiable grounding and critique alignment.
ReplicatorBench evaluates LLM agents on replicability assessment in social/behavioral sciences with incomplete data availability scenarios.
GUIDE framework resolves domain bias in GUI agents through real-time web video retrieval and annotation to improve task performance on domain-specific applications.
LiteResearcher framework addresses scaling challenges in reinforcement learning for LLM-based research agents by solving synthetic data and training cost limitations.
Safety verification method for AI agents that embeds guarantees in agentic frameworks rather than model alignment, using havoc oracle semantics.
Framework for evaluating LLM reasoning quality across six behavioral dimensions beyond final-answer correctness, providing deeper insight into reasoning processes.
Theoretical analysis showing how Bellman optimality in MDPs with catastrophic states produces prospect-theory-like behavioral signatures without utility curvature.
Survey of reasoning language model adoption across 28 scientific disciplines, identifying gaps in adoption outside hard sciences and analyzing barriers.
Analytical framework (GAMBLe) for understanding AI-Driven Research Systems coupling LLMs with automated evaluation for discovering algorithms and designs.
Study of stability-adaptivity tradeoff in hierarchical latent reasoning for long-horizon tasks, extending HRM with subgoal persistence mechanisms.
IterCAD multimodal agent framework for closed-loop CAD generation and editing through multi-turn interaction with executable sandbox, matching iterative workflows.
Theoretical analysis of logical expressiveness and structural preservation in graph neural networks through correspondence with logical formalisms.
DeXposure-Claw agentic system for DeFi regulatory supervision routing LLM decisions through forecast-grounded structured evidence to reduce false alarms.
IPO Finance Agent benchmark extending Finance Agent v2 to evaluate LLM financial analysts on IPO due diligence beyond periodic company filings.
Algorithmic techniques for teaching LLMs string matching, backtracking and error recovery to solve bit manipulation puzzles and logical reasoning tasks.
Multi-agent system for iterative ECG report generation with bidirectional editing and progressive context integration, addressing clinical workflow requirements.
Analysis of hidden inference cost from quantization of reasoning models, showing low-bit quantized models generate longer chains of thought despite correct answers.
OpenRCA 2.0 dataset with step-wise causal process labels for evaluating LLM agents on root cause analysis requiring long-context understanding and tool use.
Framework for behavioral foundation models enabling zero-shot transfer in RL by training agents to generate optimal policies for any reward function.
Algorithms for offline reinforcement learning with human feedback that are robust to data corruption from adversarial attacks or noisy preferences.
Evaluation framework assessing robustness and fairness of neural networks under adversarial perturbations simultaneously rather than in isolation.
Research on model collapse in LLMs caused by training on AI-generated content, with adaptive mitigation strategies to preserve output diversity.
Method for interpreting deep reinforcement learning policies using concept-based neuron-level analysis to improve transparency in high-stakes DRL applications.
Mantis: transformer-based foundation model pre-trained on synthetic data for time series classification using self-supervised contrastive learning.
Method for detecting LLM hallucinations in black-box settings by verifying uncertain outputs with cross-model consistency checks.
SAGE: evaluation method for LLM question-answering using search-augmented grounding instead of static references, addressing cost and reliability issues.
TraCeS method for learning safety constraints in reinforcement learning from sparse trajectory-level labels instead of dense per-timestep supervision.
Position paper arguing for interoperability standards in collaborative agentic AI ecosystems to prevent fragmentation.
Survey of AI-Generated Game Commentary systems covering multimodal perception and strategic reasoning approaches.
Evaluation framework identifying limitations in VGGSound benchmark for audio-visual foundation model assessment.
Framework for fine-tuning flow-matching models with PDE constraints to enforce physical consistency in inverse problems.
Dataset construction and multi-agent framework for training LLMs on analog circuit knowledge with structured QTSA tuples.
Comparative analysis of one-shot versus iterative pruning strategies for neural network compression.
Motion retargeting method for humanoid imitation learning from human motion data to address robotics data scarcity.
SON-GOKU scheduler using graph coloring to resolve gradient interference and partition tasks in multi-task learning.
Tool using code property graphs to detect non-local ML code smells affecting reproducibility and robustness in ML pipelines.
Agent framework for automating creation of interactive project webpages from academic papers through human-agent collaboration.
Multimodal LLM approach for GUI grounding in computer-use agents, aligning visual attention for precise screen region mapping.
Method to optimize self-consistency test-time inference for chain-of-thought reasoning in LLMs, reducing computational cost.
Research on audio-language pretraining for general-purpose audio representation, identifying barriers in scale and task coverage.
AI agent for automatically selecting optimal foundation models in remote sensing tasks with constraint awareness.
Framework for weakly supervised learning across multiple annotation patterns with theoretical guarantees for risk minimization.
HydroGym reinforcement learning benchmark platform for fluid dynamics modeling and control tasks.
Knowledge distillation technique for training smaller reasoning models by optimizing supervision allocation across prompt, chain-of-thought, and answer sections.
InfiniteWeb system for automatically generating functional web environments at scale to train GUI agents and AI assistants.
Analysis of key collision attacks on LLM semantic caching systems, demonstrating vulnerability in cache key mechanisms used by major providers.
CoReLIN framework for mobile robot navigation and manipulation in cluttered environments with persistent modifications.
ROVA framework for improving vision-language model robustness to real-world disturbances like weather and occlusion.
Online reinforcement learning technique for post-training diffusion models with variance reduction via paired trajectory sampling.