FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-Tuning
FedTreeLoRA addresses statistical and functional heterogeneity in federated LoRA fine-tuning for privacy-preserving LLM adaptation.
FedTreeLoRA addresses statistical and functional heterogeneity in federated LoRA fine-tuning for privacy-preserving LLM adaptation.
Brittlebench introduces evaluation framework measuring LLM robustness to prompt variations, typos, and real-world input noise.
Meta-research examines trustworthy AI in healthcare through interdisciplinary collaboration on transparency, robustness, and evidence standards.
Framework reframes LLMs as code generators for interpretable, verifiable decision-making in high-stakes applications with reproducible reasoning.
RelayCaching optimizes multi-agent LLM systems by reusing KV cache across agents to reduce redundant prefill computation and memory usage.
FedUAF: federated learning framework for multimodal sentiment analysis handling missing modalities and heterogeneous data distributions. arXiv paper.
Pragma-VL: method for balancing safety and helpfulness in multimodal LLMs through pragmatic arbitration, addressing jailbreaking and harmful content. arXiv paper.
FedCVR: federated learning framework for cardiovascular risk prediction using differential privacy. Privacy-preserving ML architecture case study. arXiv paper.
FRAME methodology for systematic real-world AI evaluation addressing decision-maker needs. Bridges abstract model capabilities with contextual deployment heterogeneity. arXiv paper.
ICPRL framework enabling VLMs to acquire physical intuition through interactive pixel-based reinforcement learning in dynamic environments. arXiv paper.
Hypergraph-based pre-training method for atrial fibrillation prediction in stroke patients. Machine learning for healthcare with limited data. arXiv paper.
FusionCast framework for precipitation nowcasting using asymmetric cross-modal fusion of multimodal data. Deep learning optimization for weather prediction. arXiv paper.
DreamReader: unified interpretability toolkit for text-to-image diffusion models enabling activation extraction, causal patching, ablations, and steering. arXiv research.
Research paper on safety mechanisms for diffusion and flow models, combining control barrier functions with negative guidance for constrained generation.
Empirical study of prompt-only LLM query rewriting in RAG pipelines showing domain-dependent effects on dense retrieval.
Evidence-based approach for LLMs to predict population response distributions with stability under domain shift.
Semi-supervised framework enabling single model to handle arbitrary conditional queries across heterogeneous datasets.
Self-evolving reasoning system for LLMs that prevents curriculum collapse through balanced problem space exploration.
Analysis of LLM overfitting to safety data via residual stream examination, quantifying false refusal rates in fine-tuned models.
Framework for detecting cascading risks in LLM-based multi-agent systems through semantic-geometric analysis of interaction dynamics.
Explainability method for multimodal transformers identifying synergistic and redundant cross-modal feature interactions.
Information-theoretic approach to prevent catastrophic forgetting in continual vision-language-action model adaptation for robotics.
AdaBox: grid-based clustering algorithm with adaptive hyperparameters generalizing across datasets without re-optimization.
Vision-language model fine-tuning for few-shot cross-domain learning reveals visual discriminability paradox in VLM-based tasks.
Post-training quantization applied to dataset condensation for reducing storage while maintaining model performance.
PolyGLU: transformer activation function enabling dynamic routing among multiple activation functions in feed-forward networks.
Multimodal QA framework aligning auscultation recordings with frozen LLM embeddings via gated cross-attention. Applies audio-language models to physiological signal interpretation.
FineRMoE: Architecture extending fine-grained expert design across intermediate and output dimensions in MoE models. Improves expert specialization beyond single-dimension optimization.
Decoding method using latent entropy to mitigate hallucinations in multimodal reasoning models. Extracts contextual uncertainty from token probability distributions.
Agentic LLM workflow for positioning MR spectroscopy volume-of-interest in brain tumors. Tunes placement for clinician preference and case-specific anatomy.
System for multi-version agentic analytics on lakehouses. Answers queries across competing data branches using supervaluationary semantics.
VulnAgentX: Layered agentic framework for repository-level vulnerability detection. Integrates risk screening with multi-stage validation across code structure and context.
Evaluation suite (VisualLeakBench) auditing large vision-language models for privacy vulnerabilities. Tests robustness against OCR injection and PII leakage attacks.
Experimental study comparing schema-based vs. free-form tool interfaces for LLM agents. Tests JSON Schema specifications with validation diagnostics for improved reliability.
Neuro-symbolic approach for LLM-generated formal function specifications in memory-manipulating code to enable scalable verification.
Design patterns for deploying AI agents with Model Context Protocol, identifying gaps in identity propagation, tool budgeting, and error semantics.
GPrune-LLM structured pruning method using distribution-robust neuron selection to improve cross-task generalization in compressed models.
PSKV optimization technique accelerating suffix jailbreak attacks via prefix-shared KV-cache for red-teaming LLMs.
Defense against prompt injection attacks on LLM agents using privilege separation and JSON formatting in OpenClaw platform.
OATS method for low-latency semantic routing in LLM inference gateways, selecting tools via offline embedding interpolation without GPU cost.
MIBench comprehensive benchmark for evaluating multimodal language models on various interaction patterns across modalities.
EvoClaw benchmark evaluating AI agents on continuous software evolution tasks with temporal dependencies and technical debt.
Study of dynamic sparse attention system-level challenges including cache locality, KV fragmentation, and decode throughput optimization.
Benchmark for assessing whether vision-language models can generate spatially executable robot manipulation plans from task instructions.
In-context learning method for graph foundation models enabling cross-domain alignment without modality-specific encoders.
NormCode Canvas system for LLM agentic workflows using case-based reasoning with compiler-verified scope rules to enable reliable multi-step orchestration.
Theoretical work on implementing mental-state dynamics in Transformers using biophysical principles and triadic modulation loops for attention mechanisms.
Framework reconciling in-context and in-weight learning in Transformers through dual representation space encoding.
Backdoor removal method for generative LLMs without prior knowledge of triggers or clean reference models, effective for varied tasks.
Investigation of multidimensional redundancy (spectral, temporal, spatial, semantic) in Earth observation data properties.