Learning Local Constraints for Reinforcement-Learned Content Generators
Hybrid approach combining constraint-based and RL-trained generators for game content creation with visual and global property guarantees.
Hybrid approach combining constraint-based and RL-trained generators for game content creation with visual and global property guarantees.
Causal discovery method using structural causal models and invariance principle across multiple environments.
Reinforcement learning approach to optimize qubit allocation in quantum compilers, reducing SWAP gate overhead.
Python library for conformal anomaly detection that converts anomaly scores into calibrated p-values with statistical validity.
Analysis of universal visual representations across 162 diverse vision models identifying convergent properties.
Structured pruning framework for medical SAM models preserving boundary fidelity during compression.
Secure aggregation scheme for federated learning with reduced communication rounds and dropout handling.
Fine-tuning compact LLMs with supervised learning for controllable difficulty children's story generation.
Critical analysis of 'human-in-the-loop' claims for AI safety in deployed decision systems.
Security analysis of steganographic exfiltration attacks on vector databases in RAG systems with cryptographic defenses.
Empirical comparison of dense vs sparse MoE transformers at tiny scale with parameter-matched baselines.
Theoretical analysis proving min-max optimization requires exponential query complexity for nonconvex-nonconcave functions.
Parallel scan technique for scaling recurrent neural network quantum states to larger quantum many-body systems efficiently.
Framework for iteratively evolving agentic systems via candidate generation and feedback-guided search, balancing flexibility and stability.
Study documenting failure mode where LLMs mislearn negations during fine-tuning despite recognizing false claims during training.
Verification technique for measuring input sensitivity in decision tree ensembles used for safety-critical classification tasks.
Theoretical characterization of which function classes are learnable in Valiant's original 1984 learning model with membership queries.
Benchmark framework for evaluating voice agents on realistic simulated conversations and voice-specific failure modes.
Theoretical work on generalization bounds for machine learning models on digital computers considering finite precision.
Analysis of catastrophic forgetting in LoRA fine-tuning using mean-field attention dynamics and dynamical systems theory.
Privacy analysis framework for federated learning with differential privacy providing tighter bounds across communication rounds.
Meta-learning framework for few-shot multi-task learning with limited data using linear invariant features.
Fine-tuning approach for mitigating spurious correlations caused by latent confounders during model adaptation and deployment.
Low-rank tensor approximation method for learning optimal policies in finite-horizon MDPs with high-dimensional state spaces.
Graph domain adaptation framework constructing intermediate domain sequences to handle large distribution shifts in graph learning.
Theoretical proof that transformers can exactly interpolate finite input-output sequence datasets with polynomial-sized architectures.
Regression trees algorithm for probabilistic forecasting with uncertainty quantification in critical applications.
Offline reinforcement learning method addressing value function inconsistency in model-based approaches for risk-averse applications.
Study connecting compression theory to neural network generalization, combining pruning techniques with algorithmic information theory principles.
Research on adapting State-Space Models to continual learning without stored exemplars, addressing catastrophic forgetting in evolving SSM states.
Fast two-stage approximate Top-K selection algorithm optimized for dense matrix operations on accelerators.
MaskPro enables (N:M) semi-structured sparsity for LLMs through probabilistic learning with hardware acceleration.
LoRA-Mixer routes task-specific LoRA experts through attention projections for parameter-efficient multi-task LLM adaptation.
Agent-driven mining of rare diseases from electronic health records using LLMs on noisy clinical notes.
Theoretical analysis of GPTQ quantization method for LLMs, showing connection to Babai's nearest plane algorithm.
Framework connecting foundation models to offline goal-conditioned RL, enabling test-time adaptation to novel tasks.
Analytical framework using Markov categories to explain internal mechanisms and representation learning in autoregressive language models.
Adaptive sampling approach for multi-agent reinforcement learning achieving optimal convergence in cooperative games.
Exact verification method for graph neural networks to compute adversarial robustness guarantees against attacks.
GLASS method for inference-time sparsification of LLMs using global-local importance estimation for resource-constrained deployment.
Fine-tuned vision-language model for classifying neutrino events in high-energy physics detector data.
CR-Net enables parameter-efficient LLM pre-training using cross-layer low-rank structure with reduced memory and computation.
Diffusion-based world model for offline reinforcement learning that jointly generates actions, states, and rewards.
LiLAW method that dynamically adjusts training sample weights based on difficulty for noisy data training.
Theoretical analysis of linear models for time series forecasting, examining robustness and interpretability.
Framework enabling LLMs to perform multi-step constrained reasoning by satisfying symbolic constraints in planning tasks.
Research on expert pruning vs merging for compressing Mixture-of-Experts models, showing pruning superior for generative tasks.
Latent-Augmented Discrete Diffusion: learnable auxiliary channels for improved few-step language generation.
Study of architectural factors and scaling laws trade-offs for inference-efficient LLMs.
Loopholing: mechanism preserving distributional information in discrete diffusion for parallel decoding.