A Benchmark Study of Neural Network Compression Methods for Hyperspectral Image Classification
Benchmark study comparing neural network compression techniques for hyperspectral image classification on edge devices.
Benchmark study comparing neural network compression techniques for hyperspectral image classification on edge devices.
Model Medicine: research framework for understanding, diagnosing, and treating AI model disorders through biological analogies.
Interactive Benchmarks: evaluation framework assessing model reasoning ability through active information acquisition under constraints.
CONE: hybrid transformer embedding method for preserving numerical data semantics and units in large language models.
Evaluation of GPT-5 family models on multimodal clinical reasoning tasks including diagnosis with patient data and imaging.
Method using diffusion models with contrastive signals to improve visual representations from CLIP encoders.
Machine learning research on how convolutional neural network architecture affects implicit regularization and generalization.
Speech recognition and speaker diarization system for Bengali long-form audio using Whisper and Pyannote models.
Study on data requirements and generalization of embodied foundation models for robotics and autonomous driving in interactive settings.
Machine learning research on partial label learning with instance-dependent label ambiguity, addressing instance entanglement challenges.
Research on security threats in transfer learning combined with dataset distillation, proposing osmosis distillation for model hijacking with minimal samples.
ReCouPLe framework uses natural language to mitigate causal confusion in preference-based reward learning for AI agents, improving robustness against spurious correlations.
Characterizes implicit bias of gradient descent on high-dimensional ReLU neural network regression with random features.
Visuomotor policy with working and episodic memory for non-Markovian robotic control tasks from imitation learning.
TimeWarp benchmark evaluates web agents across six UI versions spanning different internet eras using containerized environments.
Shows that replaying pre-training data during fine-tuning improves LLM performance on target domains while preventing catastrophic forgetting.
CIES metric measures robustness and stability of XAI explanations under data perturbations for business decision systems.
RepoLaunch agent automates dependency resolution, compilation, and test extraction for software repositories across languages and platforms.
GELO obfuscation scheme protects LLM prompt privacy on shared accelerators against memory access attacks.
ARC-TGI open-source framework for task-family generators that create diverse ARC-AGI puzzles with reasoning chain templates.
Novel style perturbation method for cross-domain few-shot learning addressing gradient instability and sharp minima convergence.
Theoretical analysis of analogical reasoning emergence in transformers through aligned representations and sequential training.
Optimizes vocabulary trimming in draft models for speculative decoding to balance inference latency and token coverage in LLM acceleration.
KARL system trains enterprise search agents via reinforcement learning with KARLBench evaluation suite covering six search tasks including entity search and report synthesis.
Proposes methods for training individualized decision rules with fairness constraints to mitigate discriminatory bias in personalized systems.
arXiv paper on test-time reinforcement learning to improve automatic speech recognition robustness in noisy environments and diverse accents.
arXiv paper reviewing synthetic data generation from generative AI models for statistical inference. Covers validity and reliability of synthetic data usage.
arXiv paper on ensemble methods for language models using Sequential Monte Carlo to aggregate predictions from multiple LLMs and prompting strategies.
arXiv paper analyzing chain-of-thought reasoning in models, showing performative generation where models generate tokens without revealing internal beliefs.
Method for improving masked diffusion models with path planning to enable iterative refinement during token generation.
Bi-directional adversarial network with covariate encoding for predicting remaining useful life in machinery and equipment.
Multi-scale Mamba architecture for time-series forecasting that processes inputs at multiple temporal scales using state space models.
Language agents trained to solve VM scheduling (bin packing) in cloud computing with improved generalizability and interpretability over heuristic and learning-based baselines.
VTool-R1 extends RL finetuning to vision-language models enabling multimodal reasoning with tool use and iterative image refinement during inference.
Theoretical work on attribute-efficient PAC learning of sparse halfspaces with robustness to constant-rate malicious noise in data.
CoT2 extends chain-of-thought reasoning in LLMs using continuous token representations instead of discrete sampling, with theoretical guarantees for logical reasoning tasks.
Systematic review of FPGA-based ML deployment for real-time Earth observation processing on UAVs and satellite systems with bandwidth constraints.
SPEED-RL introduces an adaptive curriculum learning method for efficient RL training of reasoning LLMs by selectively sampling intermediate-difficulty prompts to maximize learning efficiency.
Develops Bures-Wasserstein flow matching for graph generation using optimal transport-based probability paths for drug discovery and circuit design.
Proposes selective generation framework with FDR control for interactive generative systems using online learning with adversarial feedback.
Presents SKANODEs integrating structured state-space modeling with Kolmogorov-Arnold networks for interpretable discovery of nonlinear dynamical systems.
Analyzes policy network robustness using synaptic filtering to study parameter stress under internal and external adversarial perturbations.
Proposes Overtone, a dynamic patch size modulation technique for transformer-based PDE surrogate models with improved efficiency and accuracy.
Develops kernel-based maximum entropy inverse reinforcement learning for mean-field games with nonlinear reward function inference from demonstrations.
Provides theoretical analysis and improvement of GRPO, a critic-free reinforcement learning algorithm for LLM fine-tuning via group-normalized rewards.
Proposes in-training defenses against emergent misalignment in fine-tuned language models that induce harmful behaviors outside target domain.
Comprehensive survey of multi-agent reinforcement learning applications in intelligent transportation systems covering coordination and autonomous decision-making.
Introduces TabStruct metric for evaluating structural fidelity of synthetic tabular data generators against real data causal structures.
Proposes complexity-regularized proximal policy optimization replacing entropy regularization with self-regulating complexity terms for policy gradient methods.
Studies subliminal learning phenomenon where language models transfer hidden biases during distillation without explicit training signals.