torchgfn: A PyTorch GFlowNet library
Proposes pixel-level benchmark and taxonomy for VLM image tampering detection. Reformulates task with language-aware metrics beyond coarse region labels.
Proposes pixel-level benchmark and taxonomy for VLM image tampering detection. Reformulates task with language-aware metrics beyond coarse region labels.
torchgfn is a PyTorch library for generative flow networks with modular architecture. Facilitates testing new GFlowNet methods and training approaches.
Python package sbijax implementing neural simulation-based inference methods for Bayesian inference with intractable likelihoods.
Online clustering algorithm for data sequences using multi-armed bandit framework with fixed-confidence setting.
Conformal prediction method for multi-dimensional time series uncertainty quantification using flow-based models.
Bayesian optimization method with spectral mixture kernels for accelerated catalyst discovery combining categorical and continuous parameters.
Hierarchical reinforcement learning approach for scalable adaptive traffic signal control in city-scale smart city environments.
Analysis method to identify hidden breakthroughs in language model training by analyzing loss curves beyond single scalar metrics.
Predictive framework deriving scaling laws for efficient GRPO training of reasoning LLMs to optimize resource usage during fine-tuning.
Generative pipeline using 3D diffusion models to synthesize large-scale granular media assemblies in mechanically realistic configurations.
Flow-Matching Neural ODE framework for data-driven simulation of constrained multibody dynamics with supervised acceleration training.
Theoretical analysis deriving upper bounds on silhouette coefficient for clustering quality evaluation.
Theoretical and algorithmic study of weak-to-strong generalization with pseudolabels under spurious correlations from group imbalance.
Simplified graph contrastive learning method for unsupervised representation learning on heterophilic graphs without complex augmentation schemes.
Open-source Python package integrating reinforcement learning with tokamak plasma control simulators using Gymnasium environments.
Adversarial detection method for neural networks requiring no architectural changes, detecting attacks through complementary classifier label agreement.
Data annotation pipeline using LLMs with critical thinking as both annotators and judges to improve supervised learning label quality.
Efficient RL training method for reasoning LLMs using adaptive drafting to handle long-tail response generation distribution and reduce computation time.
Cross-domain offline reinforcement learning method using dynamics and value alignment for filtering datasets to improve agent training in target environments.
ReLaX: approach addressing entropy collapse in large reasoning models by promoting latent-level exploration during reinforcement learning with verifiable rewards.
Unified framework maintaining factorized momentum states across neural network training and model merging to reduce redundant computation.
Self-Distilled Reasoner: on-policy self-distillation approach for LLM reasoning that addresses distribution mismatch without teacher models.
Reinforcement Unlearning via GRPO: technique for removing sensitive data from LLMs without retraining, compliant with GDPR and EU AI Act.
Sheaf-theoretic and topological perspective on signal diffusion and attention mechanisms in graph neural networks and geometric deep learning.
Analysis of forecast uncertainty in machine learning explainability, addressing instability of LIME and SHAP near decision boundaries.
StealthRL: RL framework using group relative policy optimization to test robustness of AI-text detectors against adversarial paraphrasing attacks.
Theoretical analysis of iterative self-improvement in LLMs using reward-verified outputs with easy-to-hard curriculum learning.
Method for comparing clustering algorithms with overlapping clusters and outliers in unsupervised learning evaluation.
Spectral convolution techniques for geometric deep learning on non-Euclidean data structures like graphs and manifolds.
Interactive browser-based platform teaching federated learning concepts with real-time visualization of heterogeneous data and aggregation algorithms.
mlx-vis: GPU-accelerated dimensionality reduction library for Apple Silicon implementing 8 methods with hardware-accelerated rendering.
FEAT: linear-complexity foundation model for structured data handling heterogeneous datasets with improved attention mechanisms for large-scale applications.
Study of cone effect and modality gap in medical vision-language models, analyzing embedding concentration and cross-modal separation in supervised learning.
AcceRL: distributed asynchronous RL framework for Vision-Language-Action models with integrated trainable world models, eliminating synchronization barriers.
Difficulty-Differentiated Policy Optimization addresses Large Reasoning Models' overthinking and overconfidence by redistributing token allocation based on problem difficulty.
Framework for assessing information security awareness in LLMs, including security knowledge, attitudes, and behavior to improve rejection of unsafe requests.
Genomic language model framework using phylogenetic trees and multispecies alignment for identifying evolutionarily constrained sequences.
Analysis of multi-stage LLM inference pipelines including RAG, KV cache retrieval, routing, and reasoning with optimization strategies.
Pseudo-simulation method for evaluating autonomous vehicles addressing limitations of real-world and closed-loop simulation evaluation.
Framework for multimodal representation learning through simultaneous alignment of diverse data modalities.
Method leveraging superclasses for representation disentanglement to mitigate spurious correlations and improve group robustness.
Evaluation-Aware RL framework considers policy evaluation accuracy during training to reduce variance and bias.
Method for detecting intersectional bias in face recognition embeddings using directional alignment in latent space.
CARES lightweight module selects appropriate image resolution for vision-language models to reduce token overhead and latency.
RobotArena∞ enables scalable robot benchmarking through real-to-sim translation for evaluating diverse robotic agents.
Rep2Text framework recovers original input text from single LLM token representation using trainable adapter for interpretability.
FastMMoE accelerates multimodal LLM inference through dynamic expert activation and token pruning for reduced latency.
Unsupervised feature selection method using robust autoencoder and adaptive graph learning for high-dimensional data clustering.
Dementia-R1 applies reinforced pretraining and reasoning to LLMs for longitudinal clinical prognosis from unstructured medical notes.
Benchmark and moderation model for evaluating LLM safety, adversarial robustness, and handling of nuanced harmful content detection.