Markov Chain Decoders Overcome the Heavy-Tail Limitations of Lipschitz Generative Models
Markov Chain decoders overcome Lipschitz generative model limitations for heavy-tailed distributions. Improves VAE and Lipschitz network outputs.
Markov Chain decoders overcome Lipschitz generative model limitations for heavy-tailed distributions. Improves VAE and Lipschitz network outputs.
Hyrax: open-source modular Python framework for ML lifecycle in astronomy, supporting data acquisition through deployment on GPU infrastructure.
EgoTraj: open egocentric trajectory dataset from Meta Quest Pro for multimodal human motion prediction in robotics and navigation.
LBW-Guard: autonomous training control governance layer for stable LLM training under aggressive conditions. Operates above AdamW optimizer.
Novel conformal prediction method via transported Beta laws for calibration-conditional coverage guarantees in finite-sample settings.
RLFTSim: reinforcement learning framework for training realistic multi-agent traffic simulators via fine-tuning on real-world data.
Neuro-symbolic scenario generation using spatio-temporal logic for safety-critical testing in autonomous driving systems.
ReElicit: Bayesian optimization framework for tuning system prompts in LLMs using aggregate feedback and dynamic text embeddings.
Dataset scheduling strategy for training Audio Large Language Models that manages heterogeneity to reduce conflicting gradients and improve convergence.
GOAL: diffusion-based solver for dynamic multi-objective optimization using graph neural networks and conditioned generation.
COBALT: Cloud-based teleoperation platform for scalable robot learning via imitation using concurrent multi-user demonstrations on GPU.
Defense method against LLM backdoor attacks using rewriting with benign projections to mitigate data poisoning vulnerabilities.
ResearchArena evaluates autonomous research agents (Claude, GPT, Kimi) on their ability to conduct full research loops including ideation, experimentation, and paper writing.
Research on reducing memorization in diffusion models using Higher-Order Langevin Dynamics to address privacy and copyright concerns.
Position paper arguing mainstream uncertainty quantification in LLMs reduces to unsupervised clustering and measures internal consistency not correctness.
Framework for diagnosing multi-step reasoning failures in closed-source LLMs by attributing confidence scores to individual steps.
Physics-faithful world model for video generation preserving physical state constraints for training embodied AI systems.
Universal evaluation platform for comprehensive LLM assessment addressing limitations of static benchmarks with dynamic evaluation methods.
Study showing LLMs fail to share statistical strength across distinct presentations of same concepts in different languages/formats.
Multi-objective optimization framework for LLM agent skills under platform constraints, balancing description, instruction, and context window limits.
Controlled benchmark for measuring LLM hallucination across summarization, QA, RAG, and agentic settings with reference world models.
Neuroscience study aligning brain activity with vision-language and action models during interactive gameplay, bridging neuroscience and ML.
Training method for discrete diffusion language models using drifting objectives to improve sampling-time text generation.
Reinforcement learning environment suite for continuous-control locomotion tasks inspired by game design rather than real robotics.
Position paper analyzing Turing-completeness of autoregressive transformers and the critical role of context management in real-world systems.
Seed selection method for text-to-image diffusion models using attention dynamics on prompt core tokens to improve generation quality.
Neural surrogate planner using LSTM networks to accelerate trajectory optimization for cooperative UAV-UGV missions.
Open-source high-fidelity CFD dataset with 1800 samples for developing AI surrogate models in aircraft aerodynamics.
Deep generative models using diffusion posterior sampling for inverse problems on unstructured meshes, applied to electrical impedance tomography.
optimize_anything: LLM-based universal optimization system for text parameters achieving SOTA across six diverse domains with single model.
BCI-sift: Python toolbox for automated feature selection in brain-computer interface applications with diverse optimization algorithms.
EngiAI is a multi-agent LLM framework and benchmark for engineering design tasks combining simulation, retrieval, and manufacturing tools.
Framework extending CycloneDX standard to create AI Bills of Materials (AIBOMs) for verifiable AI system provenance and supply chain transparency.
Machine learning paper adapting conformal prediction methods to continuous AI agent evaluation with distribution-free uncertainty quantification.
Research analyzing which components of LLM agent architectures contribute most to hardware-aware code optimization, studying propose-evaluate-revise loops.
Research paper on privacy vulnerabilities in multi-tenant RAG systems when accounts collude, identifying failures in differential privacy boundaries.
PEEK system caching context orientation knowledge to improve LLM agent performance on recurring document and code repository tasks.
Novel guidance method for diffusion and flow-based generative models that preserves probability distribution geometry.
Study analyzing what mechanisms evolutionary algorithms combined with LLMs actually discover when generating and modifying code for algorithm design.
Improves Bayesian optimization by calibrating Gaussian process predictive distributions for better exploration-exploitation trade-offs.
Theoretical analysis of convergence rates for gradient descent training overparameterized neural networks with piecewise affine activations.
Federated learning approach using randomized predictors with PAC-Bayesian generalization bounds for privacy-preserving distributed training.
Faster-GCG improves efficiency of discrete token optimization attacks against aligned LLMs through better sample efficiency than prior GCG methods.
Lynx system enables efficient Mixture-of-Expert model inference through dynamic batch-aware expert selection to reduce memory bottlenecks.
Investigates how model overparameterization affects machine unlearning performance in deep neural networks through validation-based tuning.
Theoretical analysis of error propagation in score matching and diffusion-based sampling from unknown distributions in generative modeling.
Attribution method for identifying influential input regions in black-box models through minimal interpretable subset selection with combinatorial optimization.
Graph neural network architecture for interpretable drug response prediction using domain-specific prior knowledge and importance propagation.
Studies sample complexity of risk-sensitive reinforcement learning in finite MDPs using recursive entropic risk measures with a generative model.
Proposes learning abstract world models for MDPs using group-structured latent spaces with geometric priors to improve generalization from limited data.