Realistic Market Impact Modeling for Reinforcement Learning Trading Environments
Gymnasium-compatible RL trading environments with realistic nonlinear market impact models for agent evaluation.
Gymnasium-compatible RL trading environments with realistic nonlinear market impact models for agent evaluation.
Hierarchical world model with object-centric decomposition and causal latent dynamics for video prediction.
Bilevel optimization using KFAC-based hypergradients for efficient inverse Hessian-vector product computation.
Graph coarsening method for scalable Graph Convolutional Networks on large-scale node classification tasks.
Three deep learning approaches for spacecraft telemetry anomaly detection optimized for edge device deployment using neural architecture search.
Open-source Python library for machine learning on medical time-series data, addressing heterogeneous clinical data and reducing friction for ML practitioners in healthcare applications.
Lightweight uncertainty quantification for neural networks using gradient norms and isotropy assumption without training data access.
Prior-fitted tabular foundation model using in-context learning for survival analysis with limited and censored data.
Analysis showing cosine similarity between label representations in softmax classifiers does not reliably indicate model behavior.
Target-Aligned RL (TARL) framework addressing stability-recency tradeoff in target networks through selective emphasis of aligned transitions.
Graph prompt-based method for out-of-distribution detection in neural networks using disentangled representations.
Information decomposition framework measuring information spectrum in vision-language models to assess multimodal fusion vs unimodal priors.
Framework analyzing pitfalls in active learning for multimodal data, addressing missing modalities and varying interaction structures.
One-for-All: parameter-efficient LoRA variant (rsLoRA) for adapting frozen LLMs to multivariate time-series forecasting tasks.
Training-free method to combine multiple domain-specific expert LLMs into single multi-domain model without fine-tuning.
Big2Small unifies model compression techniques (pruning, quantization, distillation, decomposition) under single mathematical framework.
Uses reduced density matrices from quantum chemistry to predict phase transitions in neural networks during training and improve interpretability.
Proposes EAGLE, a federated learning algorithm ensuring fair performance across heterogeneous clients by minimizing loss gap parity.
Curvature-Guided LoRA: Parameter-efficient fine-tuning approach using prediction alignment to match full fine-tuning performance.
Study of label leakage problem in relational transfer learning where task scarcity causes models to learn task-specific shortcuts.
ShapPFN: Foundation model integrating Shapley value regression for real-time interpretable predictions on tabular data.
GPT4AP: Parameter-efficient multi-task forecasting framework using rsLoRA for air pollution prediction in data-scarce regions.
Target-Weighted Cross-Validation method for improving predictive risk estimation in spatial prediction with structured data.
Framework for discovering and validating mechanistic interpretations across neural networks to improve interpretability and generalization.
Research on Tucker attention as a generalization of approximate attention mechanisms like GQA and MLA using low-rank factorizations.
NeuralUCB-based cost-aware LLM routing algorithm that adapts online to model performance and cost, outperforming supervised routing baselines.
CRAFT: Cost-aware expert replica allocation for mixture-of-experts LLM serving with fine-grained layerwise load-balancing estimations.
Spark-LLM-Eval: Distributed framework for statistically rigorous evaluation of LLMs at scale across hundreds of thousands of samples.
UltRAG: Scalable recipe for knowledge graph RAG with LLMs to reduce hallucinations by integrating structured knowledge in context windows.
Conditional GAN for generating biomaterial microtopography with internally repeated periodic patterns and global structural consistency.
Early warning system for GPU failures using observability and structural signals beyond numeric telemetry for HPC and AI workloads.
Attention-LSTM framework for Kubernetes autoscaling addressing temporal blindness in serverless workload orchestration using deep reinforcement learning.
Foundation model for particle physics detector simulation using mixture-of-experts and parameter-efficient fine-tuning inspired by LLM techniques.
ML-based database parameter tuning system using workload compression to reduce configuration evaluation cost and improve DBMS performance.
OptiMer: Method to optimize data mixture ratios for LLM continual pre-training by extracting and merging distribution vectors without fixed hyperparameters.
Score calibration method for heterogeneous graph-vector retrieval fusion in multi-hop question answering using percentile-rank normalization.
Multi-agent reinforcement learning approach for unmanned aircraft separation assurance under adversarial GPS degradation and spoofing.
Theoretical analysis of minimum-norm interpolation under 2-uniform convexity assumptions for understanding generalization in overparameterized neural networks.
Multi-agent traffic simulation framework using self-supervised world models to scale autonomous driving system testing with unlabeled sensor data.
Model-based reinforcement learning approach using Pontryagin methods and Hamiltonian actor-critic to address compounding model errors in long-horizon value estimation.
Mimosa: evolving multi-agent framework for autonomous scientific research that synthesizes and refines LLM-based agent workflows.
PolarQuant: post-training weight quantization method for LLM compression using Hadamard rotation and Gaussian optimization.
GNN-based model for software vulnerability detection that offers better scalability than LLM approaches for code analysis tasks.
LiteCoST framework for document QA using chain-of-structured-thought and fine-tuned small language models for high accuracy and low latency.
Thiomi: large-scale multimodal dataset with 600k+ text annotations and 385k+ audio recordings across 10 African languages.
MemRerank framework distilling user purchase history into preference signals for personalized LLM-based shopping agent product reranking.
SABLE framework for semantically-aware backdoor attacks in federated learning using realistic, in-distribution triggers.
Method for generating rigorous, human-interpretable explanations for tree ensemble model predictions.
Hardware-software framework for automatic task partitioning of deep reinforcement learning on Xilinx Versal ACAP.
Fine-tuning framework (AGFT) for improving zero-shot adversarial robustness of vision-language models while preserving alignment.