Federated learning approach (DeepFusion) for training MoE-based LLMs using knowledge distillation from heterogeneous edge devices, enabling privacy-preserving distributed training.
Applies conformal Signal Temporal Logic (STL) specifications to enhance safety and robustness of RL control in aerospace (F-16 simulation), encoding control objectives formally.
Adaptive efficient rollout optimization (Train Less, Learn More) for Group Relative Policy Optimization in LLM post-training, reducing redundant rollouts when group outcomes are identical.
Zero-shot instruction following in multi-task RL using linear temporal logic (LTL) representations to specify temporally extended tasks for generalist agent policies.
Multi-class online fuzzy classifier for dynamic environments with human-defined antecedent fuzzy sets and learned consequent values in streaming data settings.
Information-theoretic framework explaining data augmentation's role in generalization and invariance learning, providing theoretical justification for augmentation effectiveness.
Framework for evaluating robustness of ML interpretability methods (LIME, SHAP) in hydrocarbon prospect risking using geophysical tabular data classification.
Addresses activation outliers in transformer quantization through spectral decay technique (S2D), establishing correlation between pre-training scale and outlier severity with theoretical analysis.
Framework analyzing reasoning modalities (code, natural language, hybrid) in LLMs under token constraints, evaluating performance tradeoffs for reasoning-specialized models.
Novel attention mechanism (SSA) replacing dot-product self-attention with Kuramoto model solution, reducing quadratic complexity and grounding in biological neural computation.
Training-free activation sparsity method (WiSparse) for efficient LLM inference considering weight-aware interactions and inter-block sensitivity, reducing computation and memory access.
Multi-agent collaboration framework for discovering latent causal variables, overcoming limitations of traditional causal discovery algorithms that assume no latent confounders.
Identifies 'silent inconsistency' in data-parallel fine-tuning of LLMs where worker-level optimization dynamics misalign despite synchronized parameters, impacting training quality.
Reinforcement learning approach (LACONIC) for controlling LLM response length during training without fixed heuristic reward shaping, addressing inference latency and computational overhead.
Multi-armed bandit algorithm (SOAR) for heterogeneous noise sources that adaptively selects data sources to minimize regret, applicable to federated or multi-source learning scenarios.
Novel parameter-efficient fine-tuning method for LLMs using mixture of experts in alternative geometric spaces (hyperbolic, spherical) to capture complex language data structures.
Theoretical analysis using numerical methods to explain why GLU variants scale better than MLPs in frontier LLMs, grounding empirical architectural choices in function approximation theory.
Open-source framework for multi-task learning to rank using transformers and self-attention for multiple relevance criteria.
Framework for multi-agent RL with dynamic agent creation and reproduction, extending MARL beyond fixed agent counts.
Semi-structured sparsity method (N:M) for training deep RL agents from scratch with hardware acceleration.
Information bottleneck regularizer for concept bottleneck models to improve interpretability while maintaining accuracy.
Lightweight adapter aligns compressed DL model embeddings with original models to improve performance in resource-constrained deployment.
Optimization method for orthogonal matrix constraints in machine learning, improving upon Landing algorithm for scalability.
SynthSAEBench toolkit for large-scale synthetic benchmarking of sparse autoencoders with realistic feature characteristics.
Systematic analysis of targeted instruction selection for LLM fine-tuning, isolating individual component contributions and best practices.
Reduces computational and memory costs in neural network training using randomized, unbiased approximations of vector-jacobian products.
D2-LoRA combines differential and directional low-rank adaptation for parameter-efficient LLM fine-tuning with algebraic mergeability.
Method to unlock latent capabilities in pretrained transformers through inner loop inference without additional training.
Theoretical framework for meta-learning defining practical universality and distinguishing algorithm-implicit learning capabilities.
Converts state-tracking tasks from sequence-to-sequence to next-token prediction format for training language models with linear RNNs.
Proposes Interactionless Inverse Reinforcement Learning to decouple alignment artifact learning from policy optimization in AI systems.
Atomix runtime providing transactional semantics for LLM agent tool use with rollback capability and side-effect safety.
Proposes curriculum learning approach (Goldilocks RL) to address sparse reward problem in RL for LLM reasoning by tuning task difficulty.
Theoretical analysis of training dynamics in reinforcement learning with verifiable rewards for transformers on compositional reasoning tasks.
Lightweight multimodal summarization framework combining web search with fine-tuned CLIP for semantic image-text alignment and summary generation.
Neural Process-based method for selecting specialized model tools in healthcare AI agents, routing queries to best-performing models per input.
Analysis of conformal prediction under distribution shift, deriving coverage guarantees for pseudo-calibrated methods using domain adaptation tools.
Theoretical analysis showing additive control variates outperform self-normalized inverse propensity scoring for off-policy evaluation in ranking systems.
Unsupervised hypergraph neural network approach addressing performance degradation on heterophilic hypergraphs without requiring labeled data.
Machine unlearning method using forget set gradients for variance-reduced (ε,δ)-unlearning with formal privacy guarantees on data removal.
Online learning approach for multi-objective prediction with adaptive algorithms handling arbitrary distribution shift and multiple simultaneous objectives.
Multimodal contrastive learning method using orthogonalization and asymmetric masking to capture modality-specific and synergistic signal interactions.
Geometric deep learning technique extending spectral convolution to orbifolds for supervised learning on non-Euclidean structured data.
New jailbreak attack method (BPJ) that automatically evades classifier-based safeguards in frontier LLMs without requiring white/grey-box access.
Theoretical analysis of discrete diffusion model sampling efficiency with sharp convergence guarantees for KL divergence using tau-leaping samplers.
First scaling law study comparing masked diffusion and uniform-state discrete diffusion language models, showing masked diffusion performance characteristics.
Novel approach to diffusion models using canonicalization instead of architectural constraints for molecular graph generation with symmetry invariance.
Benchmark study (PAPerBench) investigating how context length affects privacy leakage and personalization effectiveness in large language models at scale.
Study of geometric structures in LLM representations, showing how language statistics symmetries shape emergent geometric patterns like circular month organization and linear manifolds.
Research on source bias in dense retrieval systems, where LLM-generated text is preferentially ranked over human text due to lower perplexity in neural information retrieval.