Upper Entropy for 2-Monotone Lower Probabilities
Theoretical analysis of upper entropy computation for credal sets and uncertainty quantification. Pure mathematics focus.
Theoretical analysis of upper entropy computation for credal sets and uncertainty quantification. Pure mathematics focus.
Training method combining synthetic QA and document generation to improve LLM knowledge beyond RAG performance ceiling.
Safe reinforcement learning framework inferring constraints from user preferences with minimal expert demonstrations.
RL agent optimizing operator kernels on Huawei Ascend NPUs. Addresses knowledge gap in alternative hardware ecosystem.
Causal signal reconstruction approach for converting sparse news sentiment into reliable time series for financial/tech analysis.
StateLinFormer model using linear attention for navigation agents with long-term memory. Addresses context window limitations in Transformers.
Research on curriculum learning with dual criteria for temporal data. Proposes improved difficulty-based training scheduling.
PoiCGAN: Poisoning attack method against federated learning systems using feature-label joint perturbation.
APreQEL: Adaptive mixed precision quantization technique for deploying large language models on edge devices with reduced memory and computational requirements.
Research on how LLMs form discrete decision boundaries within continuous semantic spaces through context-driven topological distortion of number representations.
MetaKube: LLM framework for Kubernetes failure diagnosis with Episodic Pattern Memory Network that learns from operational history to improve diagnostic accuracy over time.
Theoretical framework analyzing fundamental performance limits when deploying fixed LLMs as optimization modules in agentic systems.
Method for steering code LLMs via activation space manipulation to control programming language and library preferences at inference time.
Continuous-time diffusion model for generating synthetic electronic health records with mixed numerical and categorical features.
Framework for generating interpretable explanations of learned behaviors in RL agents with formal behavior definition.
Curriculum learning approach for contextual RL using closed-form updates for self-paced task sequencing.
Lightweight fairness method for LLM-based recommenders using kernelized projection and adapters without fine-tuning.
Domain adaptation framework for foundation models using probabilistic geometric alignment and Bayesian transport.
Mechanistic analysis of grokking phenomenon in ReLU MLPs on modular arithmetic revealing algorithmic structure.
Theoretical study showing diffusion models learn manifold geometry before memorization under manifold hypothesis.
Solution for training instability in physics-informed neural networks on epidemiological models by addressing gradient pathology.
Analysis of neural collapse phenomenon in regression models across multiple layers showing low-rank structure.
Theoretical analysis revealing convex equivalences in ReLU neural networks from sparse signal processing perspective.
Kolmogorov-Arnold networks combining neural learning with symbolic structure for interpretable scientific equation discovery.
Analysis of activation function curvature role in adversarial robustness using parameterized activation family.
Study of vision-language models' robustness to distribution shifts in visual deductive reasoning tasks.
Training method for LLMs on mathematical reasoning combining RL with privileged self-distillation to improve learning on hard problems.
C++ implementation of neural network verification tool supporting bound propagation methods for DNN formal analysis.
Safe reinforcement learning method addressing constraint violations in off-policy exploration through constrained optimistic exploration Q-learning.
DIET: structured pruning method for LLMs using dimension-wise global importance scores that adapt to task-specific requirements.
Uses LLMs to generate portable patient embeddings from clinical time series that transfer across hospitals with minimal retraining.
Analysis of design challenges in iterative generative optimization using LLMs for self-improving agents; identifies hidden choices engineers must make.
Dimension-free zeroth-order estimator for PINNs addressing spatial derivative complexity and memory overhead in high-dimensional PDEs.
Iterative unsupervised framework for feature selection and clustering in high-dimensional data by recovering influential features.
Generative framework using Lagrangian relaxation-guided score-based generation to solve mixed-integer linear programming with diverse solutions.
MoE-Sieve: routing-guided LoRA fine-tuning framework for MoE models that adapts to skewed expert routing patterns for efficiency.
Investigates optimal sensor placement for GNN-based leakage detection in water distribution networks.
Dual guidance approach for RL-based LLM training combining external verification and internal experience to improve reasoning task performance.
Causal inference framework for learning disentangled representations from multiplex graphs by separating shared and layer-specific information.
RLHF-aligned LLMs exhibit response homogenization limiting uncertainty estimation; analyzes alignment tax impact across different tasks and sampling methods.
Gossip-based distributed machine learning algorithms for IoT networks with privacy constraints and limited computation/communication resources.
Graph convolutional networks using reservoir computing to address challenges with complex and dynamic graph data and long-range dependencies.
Bayesian optimization framework for tuning control policies using human preferences and pairwise comparisons instead of quantitative evaluations.
FPGA-based implementation of weightless neural networks using Tsetlin automata for on-chip training and inference with low latency and complexity.
Scalable RL pipeline for improving LLM code generation through synthetic data and curriculum learning, addressing data diversity challenges at scale.
Transformer architecture for multivariate time series forecasting using multi-resolution representations to capture short-term and long-range dependencies.
Framework uses LLMs to automatically design reward functions for cooperative multi-agent reinforcement learning, synthesizing executable reward programs from environment instrumentation.
Multi-agent reinforcement learning approach for decentralized adaptive traffic signal control using learned coordination in partially observable environments.
MolEvolve framework uses LLM guidance with evolutionary search for interpretable molecular optimization, addressing activity cliffs and lack of interpretability.
CUA-Suite dataset provides massive human-annotated continuous video demonstrations for training computer-use agents on desktop automation tasks, addressing data bottleneck.