How Embeddings Shape Graph Neural Networks: Classical vs Quantum-Oriented Node Representations
Controlled benchmark comparing classical and quantum-oriented node embeddings for graph neural network classification.
Controlled benchmark comparing classical and quantum-oriented node embeddings for graph neural network classification.
Systematic benchmark of optimizers beyond AdamW for tabular deep learning MLPs on supervised learning tasks.
BitFlipScope: framework for localizing and recovering bit-flip hardware faults in LLM parameters affecting safety.
LLMs as analytical agents detecting methodological flaws (data leakage) in published ML papers via case study.
Analysis of ICLR peer review showing large gap between score-based (91%) and text-based (83%) acceptance prediction accuracy.
Post-transformer adapter (786K params) corrects suppressed log-probabilities in aligned LLMs on politically sensitive topics.
Knowledge distillation approach enabling efficient training of State Space Models (Mamba) using pretrained Transformer models.
PolyBench: multimodal benchmark testing LLM capabilities on live prediction market data combining news and order-book dynamics.
arXiv paper on Neuro-Oracle, a retrieval-augmented agentic framework for predicting epilepsy surgical outcomes using longitudinal MRI trajectory analysis.
arXiv study analyzing Claude Code architecture and design space of agentic systems, comparing with OpenClaw, identifying five human values in agent design.
arXiv survey on explainable surrogate models for complex system simulations, examining interpretability in black-box computational models.
arXiv paper on Group Advantage Fine-Tuning (GFT) for LLMs, unifying supervised fine-tuning with reinforcement learning through policy gradient analysis.
Continual learning framework for brain disorder diagnosis from fMRI using generative replay on functional connectivity matrices.
U-Net-based deep learning model with boundary attention for glomeruli segmentation in kidney tissue using pathology foundation models.
Study of AI-assisted intervention deployment in healthcare/education with capacity constraints and imperfect user compliance.
Deep reinforcement learning for controlling rotating detonation engine mode transitions using timescale separation.
Analysis of register tokens in DINO vision transformers, showing zero-ablation overestimates their importance using multiple controls.
Analysis of synthetic data augmentation's effect on training distributions and bias-variance tradeoffs in financial machine learning.
Few-shot anomaly detection using vision-language models and heterogeneous hypergraphs for industrial and medical imaging.
Benchmark for evaluating safety of speech language models across speaker identity, acoustic style, and location contexts.
Multi-agent LLM framework using hierarchical reasoning to generate synthesizable Verilog for hardware designs, addressing context and hallucination issues.
RAG-based approach using LLMs to automate clinical value set authoring by retrieving and classifying codes from standardized vocabularies.
CURaTE: continual unlearning method for LLMs enabling real-time knowledge removal while preserving model utility.
AgentGA: genetic algorithm framework for evolving autonomous code-generation agents by optimizing agent seed.
AIPC: AI agent-driven automation system for edge model deployment targeting hardware-specific inference runtimes.
Framework for understanding mutable state layers in persistent LLM-based agents with self-modification capabilities.
RELOAD: reinforcement learning-based query optimizer for database systems with robust per-query performance.
World-Value-Action model for vision-language-action embodied agents with implicit planning capabilities.
Systematic classification and analysis of compression techniques exploiting correlations in federated learning.
Nautilus tensor compiler with automated scheduling for efficient GPU kernel generation from high-level specifications.
Bandit best-arm identification algorithm robust to both stochastic and adversarial reward distributions.
Theoretical analysis of regret tail behavior in multi-armed bandit algorithms with stochastic rewards.
arXiv paper analyzing reasoning dynamics and visual-textual information integration in 18 vision-language models.
arXiv paper evaluating multilingual text embedding models for hate speech detection across Lithuanian, Russian, and English.
arXiv paper proposing mixture-of-experts flow matching for faster language model inference while maintaining generation quality.
arXiv paper on Route to Rome Attack, demonstrating black-box adversarial suffix attacks on cost-aware LLM routers.
arXiv paper on Atropos, optimizing cost-performance trade-offs for LLM-based agents using small models with early termination and model hotswapping.
Feature selection method based on modified Shapley values for non-linear models with dependent features.
Uncertainty quantification framework for long-form LLM generation addressing factuality and coherence in open-ended text.
Machine unlearning method targeting class removal by identifying and removing forget-specific representational directions in neural networks.
Exposes vulnerability in LLM-as-judge systems where contextual framing about downstream consequences influences evaluation independent of content.
Symbolic superoptimizer for tensor programs using hierarchical symbolic graphs to represent and optimize families of implementations.
Diagnostic framework using conformal prediction and transitivity analysis to measure reliability of LLM-as-judge systems for NLG evaluation.
Controlled study examining whether LLMs can generalize systematically using shortest-path planning as a testbed to isolate training, architecture, and inference factors.
Online incremental learning method using optimal transport to manage multimodal class distributions in latent space with continuous data streams.
Survey on generative models applied to connected autonomous vehicles for predictive modeling, simulation, and decision-making.
Bilevel DPO approach for hierarchical RL addressing non-stationarity and infeasible subgoals through preference optimization.
DiffGap framework for molecule generation integrating adaptive sampling and pseudo-molecule estimation to address exposure bias in diffusion models.
Applies GNNs with human mobility data for COVID-19 forecasting, analyzing when spatio-temporal architectures outperform simpler baselines.
IMPACTX leverages XAI techniques as automated attention mechanism to improve model performance without external knowledge or manual intervention.