Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
Finch: benchmark for evaluating agents on enterprise finance workflows including data entry, retrieval, calculation, and reporting using Enron dataset.
Finch: benchmark for evaluating agents on enterprise finance workflows including data entry, retrieval, calculation, and reporting using Enron dataset.
DDFT protocol for measuring epistemic robustness in LMs under degraded information and adversarial stress beyond static benchmarks.
HAG framework for topic-adaptive agent generation in agent-based modeling balancing macro-level distributions with micro-level rationality.
ConvoLearn dataset of 2,134 tutor-student dialogues for fine-tuning LLMs on dialogic tutoring principles in science education.
Study showing LLMs exhibit robustness to emotional framing in rule-bound decision-making despite known brittleness to prompt perturbations.
TSPO: RL framework for multi-turn search-augmented LLM reasoning addressing process and reward homogenization in tool-integrated tasks.
Multi-agent LLM framework for discovering instrumental variables in causal inference through interdisciplinary knowledge synthesis.
SSLogic: agentic meta-synthesis framework where LLM agents iteratively create and refine generator-validator pairs for logic reasoning tasks.
KLong: open-source LLM agent trained for extremely long-horizon tasks using trajectory-splitting SFT and progressive RL with Research-Factory pipeline.
AI Runtime Infrastructure layer that observes and optimizes agent execution for task success, latency, token efficiency, and safety.
DeepFact benchmark and co-evolving agent system for testing factuality of search-augmented LLM-generated research reports.
HECG framework for autonomous agents using LLMs with multi-dimensional error correction and strategy transfer across tasks.
Study showing that deliberation between multiple LLMs can amplify tiny perturbations into divergent decisions, challenging robustness assumptions.
Proposes alternative training architecture for geometric and neuromorphic AI using non-standard arithmetic to reduce memory overhead.
Voxtral TTS expressive multilingual text-to-speech model generating natural speech from minimal reference audio.
ClawSafety exposes security vulnerabilities in local LLM agent frameworks where prompt injection enables privilege escalation.
AgentSocialBench evaluates privacy risks in collaborative multi-agent social networks with persistent LLM agents.
XpertBench evaluates LLM performance on expert-level open-ended tasks with rubrics-based assessment.
Addresses value hallucination in Dyna reinforcement learning agents through multistep predecessor models.
VLBiasBench evaluates biases in large vision-language models across diverse domains and question formats.
MegaFake dataset of LLM-generated fake news for understanding mechanisms behind AI-generated misinformation.
SPRIG optimizes system prompts for LLMs using genetic algorithms to improve general task performance.
Comprehensive survey of document parsing techniques for extracting structured information from unstructured documents.
RIRS framework for multi-agent RAG systems to route complex questions across distributed knowledge bases.
Human-AI collaboration for game testing using vision language models to enhance manual testing efficiency.
Reasoning Model Implicit Association Test studies implicit bias-like patterns in LLMs that use step-by-step reasoning.
BalancedDPO method aligns diffusion models with multiple conflicting evaluation metrics for text-to-image generation.
FSD bridges reasoning and decision-making in robotic manipulation by combining Vision-Language Models with action prediction for zero-shot generalization.
Bayesian ablation framework for interpreting latent task representations in neural networks, enabling probabilistic analysis of learned representations.
VERDI uses Vision-Language Models embedded in autonomous driving stack for reasoning-based trajectory planning under partial observability.
Chapter reviewing ML/AI applications in food processing, covering classification frameworks and data science approaches to food informatics.
SoSBench evaluates LLM safety alignment across six scientific domains with sophisticated, knowledge-intensive adversarial prompts.
Framework for evaluating LLM judges of LLM outputs, accounting for both sampling and judge quality uncertainty without gold-standard scores.
K-Steering enables unified multi-attribute control of LLM behavior at inference time using non-linear steering on hidden activations.
LLMs applied to combinatorial optimization of Design Structure Matrices in engineering, demonstrating reasoning capabilities for complex system reorganization.
ZINA detects and edits fine-grained hallucinations in multimodal LLMs, proposing a novel evaluation task for MLLM quality.
PRISM: lightweight fully convolutional model for multivariate time-series classification on edge devices.
Framework treating prompts as first-class citizens in LLM pipelines to enable reuse, optimization, and runtime adaptation in complex agent systems.
CATNet applies geometric deep learning (R-GCN) to catastrophe bond spread prediction in financial markets.
Embodied-R1 introduces a 3B VLM using "pointing" as unified intermediate representation to address the seeing-to-doing gap in robotic manipulation across different embodiments.
ShadowNPU system co-design for efficient on-device LLM inference on NPUs, addressing quantization sensitivity in attention operators.
DoubleAgents system for human-agent alignment in coordination tasks using a coordination agent and dashboard for preference elicitation and feedback.
Neural-MedBench reasoning-intensive benchmark for evaluating clinical reasoning ability of vision-language models beyond classification accuracy.
MedIRT psychometric framework for evaluating LLM medical competency rather than benchmark-specific performance using Item Response Theory.
ACT system combines decision trees with LLMs to provide transparent, interpretable, and auditable AI decisions on unstructured data.
Study of how autonomy levels in LLM agents affect user privacy concerns and trust, with implications for personalization design.
FURINA-Builder multi-agent pipeline for automatically constructing customizable role-playing benchmarks at scale for evaluating LLM agent behavior.
Security analysis of LLM pruning methods showing vulnerabilities in popular inference engines like vLLM when models are pruned before deployment.
Watermarking technique for LLMs using syntactic predictability to balance text quality against detection robustness for governance and trustworthiness.
XModBench benchmark measures cross-modal consistency and modality-specific biases in omni-modal large language models across audio, vision, and text.