"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
arXiv paper analyzing speech recognition failures on high-stakes short utterances. Evaluates OpenAI, Deepgram, Google, Microsoft models on U.S. street name transcription.
arXiv paper analyzing speech recognition failures on high-stakes short utterances. Evaluates OpenAI, Deepgram, Google, Microsoft models on U.S. street name transcription.
arXiv paper: tabular foundation models for association rule mining. Compares TFM performance against classical and neural ARM methods in low-data regimes.
Arbor: framework for decomposing LLM decision workflows into structured steps. Addresses instruction-following degradation in high-stakes healthcare triage scenarios.
arXiv paper: AI agents for drug asset discovery across global non-English sources. Applies wide-search agent methodology to biotech competitive intelligence and business development.
arXiv paper on energy efficiency concerns in HPC systems and AI applications. Addresses computational cost and climate impact.
Study of behavior-targeted adversarial attacks on reinforcement learning agents and defense mechanisms.
Topological methods using persistent homology to quantify semantic ambiguity in sentence-embedding neighborhoods.
Linguistic feature analysis comparing human-written and AI-generated texts across multiple dimensions.
Intermittent semi-working mask paradigm for LLMs to improve multi-turn dialogue while maintaining KV-cache efficiency.
Policy gradient theorem for Cumulative Prospect Theory objectives in reinforcement learning derived from behavioral economics.
Framework combining knowledge distillation from LVLMs with knowledge graphs for detecting toxicity in hateful memes.
NeuroLifting: graph neural network technique for efficient inference in large-scale Markov Random Fields.
Semi-supervised adversarial training via latent clustering-based data reduction to improve model robustness with less data.
Systematic evaluation of LLM ability to handle exploration-exploitation tradeoffs in contextual bandit tasks, finding reasoning models most effective.
PII-Bench: comprehensive evaluation framework with 2,842 test samples for assessing LLM privacy protection systems against PII exposure.
Neuroadaptive AI chatbot integrating real-time EEG signals to personalize learning experiences and improve engagement.
Theoretical framework for decentralized learning in biological systems interpreted as negative feedback control, with implications for AI.
Retrieval-augmented logical reasoning method to prevent LLM hallucinations caused by false premises in user queries.
Post-training quantization algorithm that iteratively rounds and updates weights while correcting errors from previous layer quantization.
Veracity Search algorithm augments chain-of-thought reasoning with correctness variables to identify errors in stepwise LLM reasoning.
New metrics for evaluating generative model quality addressing fidelity and coverage with calibration and robustness to outliers.
Discrete diffusion approach for audio inpainting using tokenized music representations to restore long missing segments.
Research on filtering pretraining data to build tamper-resistant safeguards into open-weight LLMs, addressing vulnerabilities to weight/activation modification attacks.
XAI framework using CNN occlusion maps for analyzing cough spectrograms in chronic respiratory disease diagnosis.
Out-of-distribution detection approach for continual learning in arc welding quality prediction using VQ-VAE Transformer.
Scaling long chain-of-thought reasoning in LLMs using NP-hard graph problems for cost-effective post-training dataset generation.
Evaluation of LLM-based medical assistant on realistic clinical queries beyond benchmarks, assessing contextual reasoning.
AI agents framework for automated discovery in particle physics data analysis, combining machine learning with physics knowledge.
Analysis of spectral properties in physics-informed neural networks for PDE solving, examining weight matrices and signal propagation.
Flow matching method for trajectory planning addressing long-tailed distribution in driving datasets through data balancing.
Scalable framework using local LLMs for discovering hierarchical taxonomies in large text corpora with logical partitioning.
Learning admissible heuristics for A* search with deep learning while maintaining optimality guarantees and generalization.
Self-correction framework using GPT-based vision-language models for generating reliable radiological findings in dental imaging.
Efficient linear transformation method for aligning text embedding spaces without parallel data, improving upon vec2vec.
Theoretical analysis and method for optimistic exploration bonus in reinforcement learning with human feedback.
Reinforcement learning framework (MARS-Sep) for universal sound separation using preference alignment with multimodal data.
Framework for agentic AI in software engineering beyond code, addressing socio-technical activities and research considerations.
Evaluation of small and reasoning LLMs for assessing research article quality, comparing multiple model families and techniques.
Framework synthesizing 1M+ vision-centric problems with reasoning traces and preference data for multimodal LLM training.
Consistency model approach for one-step neural decoding of error correction codes, replacing iterative diffusion decoders.
Binary Neural Networks training method using error propagation algorithm to reduce computational complexity and memory for resource-constrained devices.
Improves VAE and AE training using Random Fourier Transformation for aviation safety anomaly detection applications.
Mitigates hallucination in multimodal LLMs by exploiting full vision encoder hierarchy through text-guided layer fusion.
PROMA introduces reference-free proximal policy optimization controlling KL divergence through gradient projection for policy updates.
Proposes OPO for LLM alignment in RLHF, decoupling sampling geometry from optimization using work-dissipation principle.
Framework for improving RAG systems through embedding retrofitting and data engineering to handle annotation artifacts in text preprocessing.
Addresses path variance in score-based diffusion models, proposing MinPV principle to improve training stability and accuracy.
Meta-cognitive architecture for agentic AI in cybersecurity enabling accountable autonomous decision-making under adversarial uncertainty.
Comprehensive framework for composable skills in LLM agents, covering architecture, acquisition, security, and modularity without retraining.
Safe reinforcement learning framework using recovery-based shielding with Gaussian process models for unknown non-linear dynamical systems.