Targeted Speaker Poisoning Framework in Zero-Shot Text-to-Speech
Speech Generation Speaker Poisoning framework for removing specific speaker identities from zero-shot TTS models via machine unlearning.
Speech Generation Speaker Poisoning framework for removing specific speaker identities from zero-shot TTS models via machine unlearning.
Nwachā Munā: 5.39-hour manually transcribed Devanagari corpus for endangered Nepal Bhasha with ASR benchmark and cross-lingual transfer.
GRD-Net: generative-reconstructive-discriminative model with attention for surface defect detection and localization.
Systematic comparison of 4 training objectives for OOD detection: Cross-Entropy, Prototype, Triplet, and AP Loss.
SMAT: four-stage curriculum learning for co-adaptive exoskeleton control modeling sequential human motor adaptation.
AtomicVLA: Visual-Language-Action model for multi-step robotic manipulation with atomic skill learning and continual adaptation.
TDM-R1: reinforcement learning framework for few-step diffusion models using non-differentiable rewards like human feedback.
VoiceSHIELD-Small: lightweight real-time model for detecting malicious speech and prompt injection in voice interfaces without speech-to-text delay.
YAQIN: culturally-sensitive agentic AI application for mental healthcare support integrating Islamic frameworks for Muslim women in UK.
Ensemble approach combining RoBERTa and LLMs for aspect-based sentiment analysis in SemEval-2026 task via prediction-level fusion.
ProgAgent: continual RL agent using progress-aware reward learning from expert videos for lifelong robotic learning, addressing catastrophic forgetting.
Evaluates social and cultural biases in 7 LLMs (GPT-4o-mini, Claude, Gemini, Llama, Mistral) using Nepali cultural context with dual-metric methodology.
Method for accelerating diffusion-based text-to-image generation through pixel and timestep-level model stitching.
Theoretical advancement in temporal-difference learning for reinforcement learning addressing divergence issues in gradient TD methods.
Proposes measurement framework for addressing AI misuse in education rather than detection-based approaches to academic integrity.
Framework systematically evaluating defenses against knowledge distillation attacks on proprietary LLM APIs with taxonomy of mitigation strategies.
Open-source toolkit for steering LLMs through four control surfaces: input, structural, state, and output modifications.
Comprehensive benchmark for evaluating LLMs on complex instruction following with intricate control flows and real-world constraints.
Theoretical analysis of inference-time sampling aggregation in LLMs using particle filtering lens to understand accuracy-cost tradeoffs.
Benchmark dataset evaluating vision-language models on subtle visual differences for anomaly detection and medical imaging tasks.
Vision-only autonomy framework for robotic bronchoscopy using long-short term agent architecture for intraoperative navigation.
Test-time adaptation method using spectral experts in Vision Transformers for improved performance on distribution-shifted data.
LLM-based software agent framework using trajectory learning and entropy-aware reinforcement learning for issue fixing tasks.
Position paper proposing AI agents supervised by humans to address data-understanding imbalance across scientific disciplines.
LLM-based framework for generating human mobility trajectories during large-scale events using event-annotated datasets.
Research paper arguing that intelligence requires efficient representations and generalization rather than emergent capabilities from large model assemblies.
$OneMillion-Bench: benchmark with 400 expert-curated tasks across law, finance, healthcare evaluating language agents on real-world professional demands.
ViSA-Enhanced Aerial VLN: visual-spatial reasoning framework for aerial vision-language navigation with triple-phase collaboration.
FedMomentum: federated fine-tuning method for LLMs with LoRA that preserves training momentum via improved aggregation.
Framework reconceptualizing AI-human collaboration through alignment, process structure, and outcome quality dimensions.
MambaDance: Mamba-based diffusion model for music-synchronized dance generation emphasizing sequential and rhythmical characteristics.
DyLLM: efficient diffusion language model inference using saliency-based token selection and partial attention mechanisms.
Multimodal framework with safe cross-attention for emotion/expression recognition handling occlusions and missing modalities.
Speed3R: sparse feed-forward 3D reconstruction model reducing computational complexity from dense attention mechanisms.
ImageEdit-R1: multi-agent image editing system enhanced with reinforcement learning to handle complex multi-step user instructions.
DSH-Bench: comprehensive benchmark with hierarchical taxonomy for evaluating subject-driven text-to-image generation models.
DC-W2S method for training reliable process reward models in biological reasoning tasks using weak supervision and dual consensus.
SaiVLA-0: neuroscience-inspired vision-language-action architecture with tripartite design for compute-aware embodied AI tasks.
Foley-Flow model for coordinated video-to-audio generation using masked audio-visual alignment and conditional flows.
DARC method for inference-time LLM alignment that handles heterogeneous human preferences via risk-constrained decoding without retraining.
Framework for open-domain question answering that gradually excavates external knowledge to overcome LLM limitations in implicit reasoning tasks.
Explainable deep learning approach for fault detection and diagnosis in automotive software validation during V-development process.
Evolution strategy-based calibration method for quantizing speech processing models to address audio-specific challenges in low-bit quantization.
Research on continuous chain-of-thought reasoning in latent space for multilingual capabilities across 5 languages using GSM8k and CommonsenseQA datasets.
Analysis of speaker privacy leakage in end-to-end full-duplex speech dialogue LLM models across layers.
30B-parameter open-weight LLM for 34 European languages using curriculum learning for linguistic equity.
Multi-modal temperature and margin schedules for contrastive learning with imbalanced data.
Evaluation framework for tabular foundation models using distributional regression and proper scoring rules.
Distributed architecture enabling privacy-preserving collaboration between enterprise and cloud AI agents.
Large audio-language models for ambiguous emotion recognition with disentangled reasoning.