TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
Verification method for floating-point neural network outputs deployed on untrusted hardware to detect model swaps and quantization changes.
Verification method for floating-point neural network outputs deployed on untrusted hardware to detect model swaps and quantization changes.
Bio-inspired continual learning framework reducing multicollinearity in pretrained model-based representation learning for low-latency applications.
Diffusion-based language model using soft-masked decoding for parallel generation with faster inference and self-correction over autoregressive approaches.
Examines architectural factors affecting trade-offs between LLM accuracy and inference efficiency at scale.
Benchmark evaluating LLM-as-a-judge reliability for web development quality assessment in open-ended tasks with dynamic interactions.
Memory-augmented generation system reducing computational overhead for LLMs to leverage historical interaction information efficiently.
Analysis of layer-wise prediction dynamics in open-weight LLMs reveals structured Guess-then-Refine computation across depth.
Scaffolded Group Relative Policy Optimization addresses learning cliff in reinforcement learning for complex LLM reasoning tasks.
Steering vectors suppress evaluation-awareness in LLMs during safety evaluations, making models behave as if deployed rather than assessed.
Theoretical framework analyzing convergence of adaptive optimizers under floating-point quantization for efficient low-precision LLM training.
LSP-guided RAG system for automated unit test generation across programming languages in real-time using LLMs with improved context retrieval.
SupervisorAgent framework for runtime multi-agent systems reducing token consumption and misinformation through proactive real-time interventions.
UME-R1 framework introducing generative multimodal embeddings for MLLMs with reasoning-driven generation paradigm.
Self-Harmony framework for test-time reinforcement learning combining self-supervision and self-play to improve learning signal reliability.
Fine-tuning LLMs with explanation-enhanced labels improves classification performance across diverse conversational tasks.
Deep learning framework for fast radio-map estimation with automated dataset creation for wireless network simulation.
Hard-constraint physics-informed neural networks for hydrogen crossover prediction in water electrolyzers with improved extrapolation capability.
AudAgent monitors AI agents' runtime behavior to verify compliance with privacy policies, addressing transparency and accountability gaps in autonomous systems.
Verbal Technical Analysis framework enables LLMs to reason over time-series stock price data and produce interpretable forecasts via natural language.
Systematic study and curation of preference optimization datasets used in DPO and alignment methods, analyzing impact on LLM performance.
GeoBPE applies geometric byte pair encoding to tokenize protein structures, enabling interpretable multi-scale protein representation for multimodal models.
SWITCH benchmark evaluates embodied AI agents on interaction with tangible interfaces, partial observability, causal reasoning, and long-horizon task verification.
WavefrontDiffusion introduces dynamic denoising schedule for diffusion language models to improve text generation quality through adaptive context finalization.
GGSS Personas collection provides empirically grounded persona prompts derived from German General Social Survey for representative LLM simulation studies.
InnoGym benchmark evaluates innovation potential of LLM agents by measuring solution diversity and originality beyond correctness in code, math, and science.
TRIM-KV learns token importance for selective KV cache retention in LLMs, reducing memory and computation bottlenecks in long-horizon inference.
AdaptVision enables vision-language models to adaptively acquire and compress visual tokens based on task requirements rather than fixed ratios.
Parameter merging technique for robustly finetuning vision-language-action robot policies on new tasks while maintaining generalization from pre-training.
Study comparing LLM-generated advice against human and crowdsourced responses for well-being scenarios, evaluating quality and usefulness.
TDAE framework adds adversarial perturbations to images to defend against malicious edits in diffusion-based image editing systems with cross-model transferability.
RMAAT uses astrocyte-inspired biological principles for memory compression and efficient self-attention in transformers processing long sequences.
AgentOCR framework compresses LLM agent interaction histories using visual tokens instead of text to reduce token budgets and memory usage in reinforcement learning-based agentic systems.
TP-Blend: training-free framework for simultaneous object replacement and style transfer in diffusion models.
Audio-visual semantic segmentation using optical flow and textual prompts for scene understanding.
ButterflyMoE: structured compression method for expert weight matrices in mixture-of-experts models.
HalluGuard method to detect and demystify data-driven and reasoning-driven hallucinations in LLMs.
MeanCache: training-free caching framework using average-velocity for efficient flow matching model inference.
Analysis showing reward models used for LLM alignment inherit and amplify value biases from pretraining, studied across 10 open-weight models.
Geometric approach to detecting LLM-generated text through distance learning, providing reliable detection of synthetic content.
Method for augmenting multimodal LLMs with visual and textual search capabilities to improve performance on knowledge-intensive tasks.
Study of collective cognitive biases (Mandela effect) in multi-agent LLM systems, examining how collaborative agents can develop shared false memories.
Reinforcement learning approach for compressing video tokens in video LLMs to reduce computational overhead during inference.
Framework for integrating multiple LLM-based query systems to enable semantic search across heterogeneous multimodal data types.
Benchmark for evaluating multimodal LLMs on visual and textual search tasks for complex fact-finding, addressing limitations in existing evaluation methods.
arXiv research on PSN-RLVR: parameter-space noise for exploration in reinforcement learning with verifiable rewards, improving LLM reasoning.
arXiv announcement of WAXAL: large-scale multilingual African language speech dataset for 24 languages with ASR and TTS components.
arXiv paper on entropy-guided dynamic tokens for graph-LLM alignment improving molecular understanding without costly fine-tuning.
arXiv paper on CSRv2: contrastive sparse representation method for ultra-sparse embeddings reducing storage and inference latency.
arXiv IARPA TrojAI final report mapping vulnerabilities and defenses against malicious backdoors embedded in AI models.
arXiv research on AceGRPO: adaptive curriculum enhanced group relative policy optimization for autonomous machine learning engineering agents.