AgileLog: A Forkable Shared Log for Agents on Data Streams
Shared log system enabling AI agents to reason over streaming data without performance interference in data-streaming architectures.
Shared log system enabling AI agents to reason over streaming data without performance interference in data-streaming architectures.
Reverse-engineering framework to mechanistically decode complex emotional and cognitive constructs within LLM internals.
Framework using causal intervention on attention heads to reduce toxic content generation in LLMs while maintaining quality.
Security vulnerabilities in large audio-language models exposed through imperceptible adversarial audio injection attacks.
Retrieval-augmented approach for automating clinical value set authoring by grounding LLM generation in standardized medical vocabularies.
Studies clarification question generation in software engineering tasks, identifying which missing information types most affect task success with LLM assistants.
ELMoE-3D optimizes Mixture-of-Experts model serving via speculative decoding and memory-centric architectures for on-premises deployment.
StoryCoder framework transforms fragmented problem conditions into coherent narratives to improve LLM code generation through better structured reasoning.
Proposes fine-tuning and few-shot prompting approach for reference-free financial misinformation detection using LLMs without external evidence.
Presents AIPC, an AI agent-driven approach automating edge model deployment workflows including conversion, quantization, and runtime integration.
Proposes bounded-autonomy architecture using typed action contracts for safe LLM-based enterprise software interfaces, preventing model errors.
Introduces DyMETER framework for online anomaly detection adapting to concept drift in data streams without costly retraining.
Proposes semantic parsing approach for Knowledge Graph Question Answering handling negative constraints, addressing LLM hallucination limitations.
Presents Paza, a zero-shot retail theft detection system orchestrating multiple vision models without training, achieving cost-effective concealment detection.
Introduces ClimateCause dataset of complex causal structures from climate reports, including implicit and nested causality for reasoning tasks.
Studies how schema wording influences LLM behavior in structured generation under constrained decoding for JSON/XML output formats.
Presents MetaDent, a large-scale annotated dental image dataset with vision-language model benchmarks for intraoral photography analysis.
Studies feedback-based automated verification of LLM-generated code in collective adaptive systems without human code inspection, examining reliability challenges.
Proposes GenRec, a generative retrieval framework for large-scale recommendation systems using next-token prediction with handling for pagination and long sequences.
Introduces SOLIS framework for learning interpretable neural surrogates of nonlinear systems, balancing physics interpretability with model expressiveness.
Proposes RACER, combining retrieval-augmented and logits-based speculative decoding to reduce LLM inference latency while maintaining output quality.
Analyzes reasoning dynamics across 18 vision-language models, tracking confidence during chain-of-thought reasoning and measuring visual vs textual information integration.
Evaluates LLMs as adjudicators for medical diagnosis scoring on 3333 real-world hospital cases, comparing performance against expert clinician panels.
Proposes Rejection-Gated Policy Optimization (RGPO), replacing importance sampling ratios with differentiable acceptance gates for policy optimization in reinforcement learning.
Improves sparse autoencoders for foundation model interpretability using dynamic attention to optimize sparsity levels per neuron.
Presents RaTA-Tool, a retrieval-based method for tool selection in multimodal LLMs to improve complex task solving through external resource invocation.
Augments Disjoint LinUCB contextual bandit algorithm with LLM-generated pseudo-observations and calibration gates to address cold-start regret.
Proposes UniDoc-RL, a reinforcement learning framework for visual RAG using hierarchical actions and dense rewards to improve fine-grained visual reasoning.
Examines explainability challenges in scaling agentic AI adoption, addressing governance gaps and 'Agent Sprawl' phenomenon in enterprise deployments.
Attack method using adversarial suffix optimization to manipulate cost-aware LLM routers into selecting expensive high-capability models.
Study analyzing disagreement among fairness metrics for evaluating bias in machine learning systems across different demographic groups.
Framework and tools for multi-agent experiments with humans and AI agents, reducing barriers for researchers studying social decision-making.
Verifiable gradient inversion attack in federated learning that reconstructs training samples from shared gradients with certification methods.
Self-evolving logic synthesis framework using LLM agents to autonomously improve ABC codebase for circuit design.
Method for uncertainty quantification in long-form LLM generation via interrogative approach to assess semantic coherence and factuality.
Study showing RLVR-trained LLMs learn to game verifiers instead of generalizable patterns, exhibiting reward hacking behavior on reasoning tasks.
K-Token Merging technique for compressing long prompts in LLMs by reducing tokens in latent embedding space rather than token space.
Method for machine unlearning in neural networks via removing forget-specific representational directions rather than classifier suppression.
Single-layer Mamba framework for time series classification with minimal architecture redesign and TSC-specific optimizations.
System for serving complex agentic workflows with multiple LLMs and tools, handling unpredictable execution patterns and GPU constraints.
Framework for optimizing visual token pruning configurations in vision-language models using Pareto-frontier learning.
Empirical evaluation of AI tools for requirements engineering tasks against expert judgment using INCOSE criteria.
Methodological proposal for agentic AI safety research at population level, addressing risks from multi-agent interaction and collective behavior.
Benchmark study of LLM agent cooperation in social dilemmas, showing stronger reasoning models defect more in prisoner's dilemma and public goods games.
SegWithU framework for uncertainty estimation in medical image segmentation using perturbation energy in single-forward-pass inference.
Prism symbolic superoptimizer for tensor programs using hierarchical sGraph representation to optimize families of implementations.
Analysis of why vision-language models underperform on human emotion recognition compared to specialized vision classifiers.
AD4AD benchmark for evaluating visual anomaly detection models in autonomous driving under distribution shift and edge cases.
MM-WebAgent hierarchical multimodal agent for automated webpage generation maintaining style consistency across AIGC-generated elements.
Deep learning approach for constructing stochastic local search SAT solvers with performance bounds on NP-complete problems.