Agentic Publications: LLM-driven framework transforming scientific papers into interactive knowledge systems integrating structured data and multimedia.
Position paper arguing tool-augmented agents should invoke external tools only when epistemically necessary, not treating tools as ordinary actions.
VCBench: First benchmark for predicting founder success in venture capital using LLMs, with sparse signals and uncertain outcomes exceeding market index performance.
Combines LLM agents with operations research algorithms for inventory control, enabling flexible reasoning while adapting to demand distribution shifts.
HiMAC: Hierarchical macro-micro learning framework enabling LLM agents to handle long-horizon tasks with structured planning and reliable execution via separate reasoning/action generation.
Framework for quantifying trust in autonomous AI agents through end-to-end operational outcomes rather than model-internal properties, relevant to deployed agents with financial risk.
Soft Tournament Equilibrium: Framework for evaluating non-transitive LLM-based agents using set-valued cores instead of linear rankings for cyclic competitive domains.
HiL-Bench evaluates whether coding agents know when to ask for help versus act autonomously on incomplete specifications, addressing judgment gaps in frontier agents.
HWE-Bench: First large-scale benchmark with 417 real hardware bug repair tasks for evaluating LLM agents at repository-level, beyond component-level HDL generation.
Mechanized proofs in Coq for structural governance of cognitive workflows using coinductive logic and the Interaction Trees library.
On-policy self-distillation for GUI grounding in autonomous agents, improving performance with dense signals without expensive rollouts.
Brain-inspired continual learning algorithm for spiking neural networks that adaptively reorganizes neural pathways across multiple tasks.
Deep deterministic policy gradient with symmetric data augmentation for sample-efficient offline reinforcement learning in aircraft control.
Equivalence between collective decision-making and reinforcement learning through analysis of honey bee nest-hunting behavior.
Study on bias inheritance in LLM-based synthetic data augmentation, showing how models amplify training data biases in downstream tasks.
AhaRobot: low-cost open-source bimanual mobile manipulator for embodied AI and vision-language-action model training.
GPU-friendly algorithms for polar decomposition and matrix sign functions applied to the Muon optimizer for deep neural network training.
Selective Jacobi decoding to accelerate inference in discrete autoregressive normalizing flows, improving speed of generative models.
Scaling Bayesian optimization to many observations by replacing Gaussian process surrogates with more scalable alternatives.
RoboEval: structured evaluation framework and benchmark for robotic manipulation with behavioral and outcome metrics beyond binary success.
ReCode: reinforcement learning method for code generation using reasoning-process rewards to improve correctness and reasoning quality.
SPRINT: fingerprinting technique to detect source models of AI-generated images and defend against adaptive attacks on model attribution.
LinkAnchor: autonomous LLM-based agent for recovering issue-to-commit links in GitHub repositories to improve software traceability.
Approach using multi-language models for real-time syntax highlighting in web-based development tools with strict latency and memory constraints.
Empirical study showing equivariant architectures demonstrate better power-law scaling than non-equivariant models for learning interatomic potentials.
Memory-augmented framework enabling LLM agents to learn from labeled examples without parameter updates using episodic and semantic memory.
Theoretical framework proposing semantic information theory for LLMs as alternative to bit-based information theory, addressing foundational understanding.
System for efficient multi-turn LLM agent scheduling using KV cache time-to-live to enable cache reuse across interleaved tool calls.
Analysis of how LLMs internally structure memorized training data using citation generation as proxy, examining hierarchical memorization patterns.
Study evaluating LLM capability for personalized access control decisions in applications and agent-based systems to reduce user cognitive burden.
RL post-training approach using bootstrapped mixed rewards and canonical action order hints to improve Transformer performance on structured problems.
Benchmarking framework for evaluating PDF document parsers on mathematical formula extraction using synthetically generated test PDFs with LaTeX ground truth.
Open-source federated platform for operationalizing AI/ML lifecycle in scientific research with FAIR principles and MLOps tooling.
Framework for intelligent knowledge mining from disparate data sources using AI analysis with emphasis on trustworthy data preservation and integration.
Case study comparing LLM performance to mental health professionals on personality disorder diagnosis from Polish-language patient narratives.
Multimodal framework combining meteorological text with visual data for short-term precipitation nowcasting using language-aware constraints.
Benchmark dataset (LitVISTA) for evaluating LLM narrative structure and story arcs in literary text generation versus human narratives.
LLM-supported framework for translating conceptual queueing system descriptions into verified executable simulation programs with mechanism fidelity checking.
Benchmark for tree canopy segmentation from aerial imagery using fine-tuned pretrained models on small dataset of 150 annotated images.
Privacy evaluation of multimodal RAG systems through membership inference and caption retrieval attacks on vision-centric tasks.
Research on scaling cooperative multi-agent reinforcement learning by addressing cross-agent noise in shared reward signals, relevant to differentiable systems.
Analysis showing test-time training with KV binding is equivalent to learned linear attention, reinterpreting online meta-learning.
Sparse attention mechanism using online permutation and early stopping to improve long-context LLM inference efficiency.
Benchmark with 100 web app specifications and 10K+ substeps for evaluating AI code generation on end-to-end development tasks.
Multi-granularity credit assignment method for reinforcement learning on multi-turn emotional support dialogue with sparse rewards.
Hierarchical specialization approach for respiratory audio question answering in healthcare conversational AI.
Self-improvement framework maximizing mutual information between prompts and responses for LLM personalization without additional labeled data.
Method for improving LLM factuality evaluation robustness by reducing candidate-order sensitivity through permutation consensus.
Modality-aware scheduler for multimodal LLM inference reducing latency and memory contention from heterogeneous workloads.
Single-agent robotic architecture organizing intelligence through embodied capability modules with unified identity and control.