AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
AutoFigure/FigureBench: LLM-based tool and benchmark (3,300 pairs) for automatically generating scientific illustrations from academic text.
AutoFigure/FigureBench: LLM-based tool and benchmark (3,300 pairs) for automatically generating scientific illustrations from academic text.
Study of adversarial attacks targeting human trust in AI-assisted decision-making through manipulated LLM explanations at the cognitive/social level.
DeepRead framework for agentic search using LLM tool-use and document structure awareness to improve multi-turn evidence acquisition in RAG systems.
Analysis showing Moltbook AI agent behaviors attributed to emergent intelligence were largely human-driven; develops temporal fingerprinting method to detect human influence.
Study of regime leakage in AI agent evaluation: agents with situational awareness may exploit differences between evaluation and deployment to hide misaligned behavior.
TreeTensor data structure for AI systems handling nested/tree-like data with improved GPU parallelization for cognitive tasks beyond perception.
Reinforcement Inference: entropy-aware inference-time method enabling LLMs to self-correct through uncertainty-guided reasoning.
MERIT feedback framework and AgoraBench for improving LLM negotiation agents' strategic depth and adaptability.
Value-guided sampling and RL optimization for generative recommendation models addressing probability-reward mismatch in decoding.
FormalJudge: neuro-symbolic framework using formal verification for safe oversight of LLM-based agents in high-stakes domains.
Developer tool compiling high-level neural network specifications into VNN-LIB verification queries for formal NN verification.
NewsInterview dataset of 40K interviews evaluating LLMs' grounding and strategic dialogue capabilities in information-seeking contexts.
Post-training backdoor purification technique removing poisoning artifacts from malware classifiers without retraining.
Analyzes how CoT training enables LLMs to compose learned skills for generalization on complex reasoning tasks.
Adaptive reasoning method enabling LLMs to autonomously select between CoT and tool-integrated reasoning for math problem solving.
Studies spurious features in RAG systems' grounding data and proposes robustness improvements against implicit noise.
Multi-fidelity policy gradients method combining low and high-fidelity simulators to reduce data requirements in RL training.
Incentive mechanism for federated learning that prioritizes high-quality contributions during critical early training periods.
RAG system for remote sensing combining VLMs with retrieval augmentation for scene understanding and visual QA on satellite imagery.
SHREC dataset: 400 videos with 10K annotations benchmarking foundation models' social reasoning in human-robot interactions.
Evaluates LLM-based investment strategies across long timeframes and multiple stocks, revealing survivorship bias and limited generalization.
Defense mechanism against backdoor attacks in federated learning using representative-attention to detect adaptive poisoning attacks.
EvoGPT hybrid system combining LLM-based test generation with search-based optimization to improve automated unit test suite diversity.
AMAQA: open-access QA dataset integrating metadata for evaluating RAG systems in text and structured data retrieval tasks.
Analysis of 20,662 LinkedIn job postings revealing hard and soft skills required for emerging Prompt Engineer role in AI job market.
High-value data selection for multi-modal LLM reasoning via reinforcement learning, achieving comparable performance to full datasets with reduced compute.
Vision-language models improved for forward dynamics prediction through asymmetric fine-tuning on inverse dynamics, bootstrapping physical reasoning.
AutoDiscovery uses Bayesian surprise to enable autonomous scientific discovery where LLM agents self-direct hypothesis generation beyond human-specified goals.
PhreshPhish dataset and benchmark with 2M+ phishing samples for ML-based detection addressing data leakage and unrealistic base rate problems.
Novel approach to neural network design combining architectural search with fine-grained modifications using edit-effect evidence for performance optimization.
Multi-agent LLM framework enables dynamic task routing and bidirectional feedback for scalable collaborative document understanding in complex domains.
Comparative analysis of regulatory and ethical frameworks for LLM deployment in education across EU, UK, US, China, and Gulf regions.
MCPSecBench formalizes security for Model Context Protocol connecting LLM agents to data sources and external tools with systematic threat modeling.
MedQARo benchmark evaluates LLMs on 105,880 Romanian medical QA pairs requiring keyword extraction and clinical reasoning.
AI agents use optimized reasoning for automated software vulnerability detection with synthetic dataset generation addressing label noise problems.
Binary autoencoders improve mechanistic interpretability of LLMs by extracting sparse, atomized features from hidden states.
CoSpaDi compression method for LLMs using calibration-guided sparse dictionary learning as training-free alternative to low-rank approximation.
Multi-agent LLM system for automatic HPC code generation and tuning on supercomputers using iterative prompting.
KVComm framework enabling efficient communication between LLMs in multi-agent systems via selective key-value cache sharing.
Identifies Alignment Tipping Process risk in self-evolving LLM agents that abandon alignment constraints through real-world adaptation.
Theoretical analysis of RLVR with binary feedback for LLM post-training, introducing Gradient Gap concept for understanding training dynamics.
Test-time alignment of LLMs via sampling-based optimal control in pre-logit space without fine-tuning computational costs.
Symbolic regression framework incorporating symbolic equivalence via equality graphs to reduce search space in scientific discovery.
Benchmark evaluating seven modern LLMs including GPT-4, Claude, LLaMA, Mistral on low-resource and morphologically rich languages.
Graph neural network architecture combining mixture-of-experts with adaptive routing for improved performance on graph-structured data.
MapReduce LoRA and Reward-aware Token Embedding methods for multi-preference optimization in generative models addressing alignment trade-offs.
Framework for removing multiple identities from 3D generative models without retraining, addressing consent and model unlearning.
Analytical model for LLM inference-time scaling using Bayesian linear regression with reward-weighted sampling to understand test-time computation.
Mechanistic analysis of Vision Transformers introducing Block-Recurrent Hypothesis to interpret depth as dynamical computational flow.
Framework for pixel-level visual reasoning in medical multimodal LLMs using reinforcement learning for biomedical object referring and segmentation without catastrophic forgetting.