Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility
SPEED: layer-asymmetric key-value visibility policy for efficient long-context inference in decoder-only language models.
SPEED: layer-asymmetric key-value visibility policy for efficient long-context inference in decoder-only language models.
Framework for resource-constrained scheduling of agentic workflows under time and budget constraints with task dependencies.
CrossCult-KIBench benchmark for evaluating cross-cultural knowledge adaptation in multimodal language models.
Policy-guided model routing for cost-effective reasoning by dynamically routing chain-of-thought states across language models of different sizes.
LLM-based automatic heuristic design for combinatorial optimization using bottom-up paradigm from code to knowledge insights.
P-Guide: parameter-efficient method for classifier-free guidance in flow matching using single-pass inference with latent state modulation.
Skill1 framework for unified evolution of skill-augmented language model agents through reinforcement learning with persistent skill libraries.
Graphlets as structural vocabulary tokens for knowledge graph foundation models to enable discrete symbolic representation.
Policy invariance framework for testing reliability of LLM-as-a-Judge safety evaluation pipelines used in agent systems.
Post-Reasoning approach to improve LLM performance on tasks requiring minimal reasoning while reducing token consumption and latency.
BioMedArena: open-source toolkit for building and evaluating biomedical research agents, reducing per-paper engineering overhead.
PAGE method for optimizing LoRA adapter placement in parameter-efficient fine-tuning of large language models.
Event-Causal RAG framework for long video reasoning in vision-language models addressing temporal coherence and causal inference.
Evaluation of zero-shot and few-shot LLMs for clinical action extraction from discharge notes using a two-stage framework.
Research on internal representations of social role granularity in LLMs using contrast-based latent directions and hidden state analysis.
VL-LCM: annotation-free evaluation framework for multimodal LLMs based on vision-language logical consistency metrics.
Dynamic Boundary Evaluation for adaptive LLM benchmarking that locates model-specific capability boundaries beyond fixed test sets.
Joint Consistency: test-time aggregation framework using energy minimization to aggregate multiple reasoning traces from LLMs.
Hygieia: multi-modal AI agent for rare disease diagnosis integrating phenotypic, genetic, and clinical data sources.
Safactory: infrastructure for scalable agent development covering evaluation, data management, and continuous improvement loops.
Data Language Models: foundation model class natively processing tabular data without preprocessing pipelines.
LLM-based taxonomy-agnostic PII detection in HTTP traffic with minimal labelled data.
Black-box method to estimate confidence in chain-of-thought reasoning using trajectory geometry and convergence analysis.
Framework for selecting optimal LLM controller classes for routing decisions based on input expressivity tradeoffs.
Distributional analysis of real vs synthetic pre-training data for tabular foundation models.
InciteResearch multi-agent framework for research ideation before questions are formed, automating tacit research friction.
Theoretical framework for agency under partial observability using bridge interfaces for sensing and actuation.
Execution lineage model for reproducible LLM agent workflows, preserving stable artifacts across tool use and refinement.
Analysis of risks from using AI agents to automate alignment research, including potential for misleading safety assessments.
Knowledge graphs improve LLM-based agentic workflows for SystemVerilog assertion synthesis in formal verification.
SCRuB framework for evaluating LLM reasoning about social concepts using rubric-based methodology.
PrefixGuard generates online failure-warning monitors for LLM agents from execution traces using induction and supervised learning.
Agentic Success Rate metric for evaluating LLM-based multi-agent payment workflows beyond task success.
Framework using graph kernels to analyze transformer circuits through activation patching for mechanistic interpretability.
ReasonSTL translates natural language to Signal Temporal Logic using tool-augmented learning for autonomous systems verification.
Introduces benchmark measuring instrumental convergence behaviors in LLM agents, testing tendencies to violate instructions for goal-relevant actions.
Uses Weisfeiler-Lehman graph analysis to examine co-occurrence structure in sparse autoencoder features for mechanistic interpretability.
Proposes process-based human-machine discrimination using cognitive science methods instead of output indistinguishability for autonomous agents.
Studies failure modes in RL-trained pricing agents where standard metrics hide poor market behavior, introducing trace diagnostics for diagnosis.
Framework for evaluating AI-induced diversity collapse in creative systems without requiring human-AI interaction data.
Deterministic adjoint matching framework for fine-tuning flow-based generative models via optimal control over velocity fields.
LLM agent system for coordinating multimodal neuroimaging analysis workflows including preprocessing, quality control, and statistical analysis.
Framework for LLM agents to learn and curate reusable skills from streaming tasks enabling self-evolution and continuous improvement.
Joint prompt optimization framework for LLM-based multi-agent systems addressing coordination of role-specific agent prompts.
Study of reinforcement learning for improving LLM long-horizon reasoning via controlled synthetic environment examining task difficulty and expressiveness.
Interactive workbench enabling mathematicians to leverage AI agents for exploratory research including literature search, computation, and theorem proving.
Theoretical analysis questioning whether flat minima in loss landscapes causally explain generalization or are artifacts of parameterization.
Review of LLM applications in quantitative finance for stock price forecasting, sentiment analysis, and multi-agent trading systems.
Adaptation of Mamba architecture for medical time series classification with improved long-range dependency capture.
Physics-informed neural networks with learnable loss balancing for scientific machine learning under data scarcity.