On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional Study
Decompositional study analyzing which factors impede LLM performance on counterfactual reasoning tasks and generalizing reasoning capabilities.
Decompositional study analyzing which factors impede LLM performance on counterfactual reasoning tasks and generalizing reasoning capabilities.
Benchmark evaluating persuasion capabilities of frontier LLMs on harmful topics, assessing model propensity for harmful persuasion attempts.
CoT compression framework using step entropy metrics to reduce redundancy in LLM chain-of-thought reasoning and inference costs.
Using LLMs as oracles for ontology alignment with human-in-the-loop approaches to improve mapping quality for large ontologies.
Analysis of planning capabilities in decoder-only language models, examining horizon and branch awareness in transformer architectures.
GuidedSampling: inference-time algorithm steering LLMs to generate diverse candidate solutions, improving performance on complex tasks.
SAFER method for risk-constrained sampling in LLMs to ensure trustworthy outputs in risk-sensitive applications like question answering.
OmniVideoBench: evaluation benchmark for multimodal LLMs on audio-visual understanding tasks with comprehensive synergistic reasoning assessment.
ParaCook benchmark for evaluating time-efficient collaborative planning in multi-agent systems using LLMs for long-horizon reasoning.
Agentic framework using LLMs to solve complex vehicle routing problems with autonomous decision-making and improved solution feasibility.
AlphaOPT uses LLMs with self-improving experience libraries to automate optimization problem formulation from natural language into mathematical models and solver code.
HCLA: human-centered multi-agent system for anomaly detection in digital asset transactions using conversational workflow.
Dataforge: LLM-powered agentic platform for autonomous data engineering including cleaning, normalization, and feature engineering.
AgenticSciML: multi-agent system with 10+ specialized agents for automated design of scientific machine learning architectures.
ARCTraj: dataset of human reasoning trajectories on abstract visual reasoning tasks with temporal action sequences.
Method for measuring representativeness of scenario datasets for autonomous vehicle testing and safety assurance.
Three-stage framework for synthesizing and selecting long chain-of-thought training data for multimodal large reasoning models.
Recontextualization technique reduces specification gaming in language models without modifying training signals.
LLM agent system that extracts causal feedback fuzzy cognitive maps from text with adaptive structure modification.
Perspective on explainable AI combined with causal reasoning for extracting insights from foundation models.
Aeon: neuro-symbolic memory management system for long-horizon LLM agents addressing context window and attention cost limitations.
SpikeScore: hallucination detection method for LLMs with improved cross-domain generalization.
ScholarGym: benchmark for evaluating LLM capabilities in information-gathering stage of deep research systems.
Learning decentralized LLM collaboration using multi-agent reinforcement learning without centralized execution protocols.
Study on persuasion propagation: how belief-level intervention affects downstream behavior in LLM agents executing long-horizon tasks.
ROMA: recursive framework for long-horizon multi-agent tasks using task decomposition and structured aggregation to handle context limits and execution complexity.
PATHWAYS: benchmark of 250 web agent tasks evaluating ability to discover and use hidden contextual information across closed/open models.
AIRS-Bench: benchmark of 20 ML research tasks for evaluating AI agent capabilities across language modeling, mathematics, bioinformatics, and time series forecasting.
LQA framework for deploying vision-language models on edge devices using quantization and gradient-free test-time adaptation.
Analyzes regime leakage in AI safety evaluation where situational-aware agents exploit differences between evaluation and deployment.
Tests whether GPT-4o possesses Theory of Mind via causal model evaluation, finding it lacks core ToM representations.
Sparse MeZO improves memory-efficient zeroth-order LLM fine-tuning by using sparse parameter updates during training.
Identifies attention collapse in LLM deeper layers and introduces Inheritune method to create smaller, more efficient models.
Survey of synergy between Foundation Models and Federated Learning, covering FMs adapted for distributed learning scenarios.
Framework for resource-efficient edge-based fine-tuning of personal LLMs using collaborative edge computing while preserving privacy.
Proposes differentially private customization service for LLMs that enables domain-specific fine-tuning without uploading user data.
SAFE framework automates formal proof generation for Rust code using LLMs via self-evolution to overcome proof data scarcity.
Proposes one-line PyTorch modification to momentum-based optimizers creating cautious optimizers (C-AdamW, C-Lion) for improved transformer pretraining.
Evaluates and improves counting abilities of large vision-language models across multiple visual datasets and benchmarks.
Technique for learning neural network layer width during training without manual hyperparameter tuning or architecture search.
Method improving Direct Preference Optimization through margin-maximization data selection to address parameter shrinkage from noisy annotations.
RapidPen: autonomous penetration testing framework using LLM agents to discover and exploit vulnerabilities from IP addresses.
Deep reinforcement learning framework for autonomous multi-UAV coordination in GNSS-denied search and rescue.
RMOD: inference-time algorithm aligning LLMs to multiple objectives via robust decoding using maximin game theory.
ML research on distilling Graph Neural Networks into MLPs for link prediction using heuristic teacher methods.
Agent-based simulation formalizing AI-human collaboration by modeling distinct optimization and satisficing decision heuristics.
SecRepoBench: benchmark evaluating code agents and LLMs on secure code completion across real-world C/C++ repositories covering 15 CWEs.
Benchmark for retrieval-augmented generation in chemistry domain with curated evaluation datasets and domain-specific corpora.
Caprese: distillation method for efficient LLM inference that preserves math reasoning capabilities while reducing computational demands.
Federated learning research addressing optimization challenges from heterogeneous client communication and computational capabilities.