Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
Inference-time distillation enables cost-efficient LLM agents without fine-tuning or prompt engineering, preserving iteration agility for batch jobs.
Inference-time distillation enables cost-efficient LLM agents without fine-tuning or prompt engineering, preserving iteration agility for batch jobs.
Joint distillation technique accelerating likelihood evaluation and sampling in flow-based generative models by reducing neural function evaluations.
DFedReweighting: Framework for objective-oriented reweighting in decentralized federated learning addressing fairness and Byzantine robustness challenges.
Torch Geometric Pool: PyTorch library standardizing graph pooling methods with unified SRCL interface for easier comparison and reuse across different pooling approaches.
Unified framework connecting discrete, Gaussian, and simplicial diffusion models for sequences like DNA, proteins, and language.
Incomplete multi-view clustering method using Missing Pattern Trees to improve pair utilization across inconsistent missing patterns.
Multi-agent option discovery method using inter-agent relative representations to improve coordination and reduce joint state space complexity.
Forest proximity computation via Separable Weighted Leaf-Collision kernels for improved scalability in kernel and representation-learning pipelines.
FOREVER: Forgetting curve-inspired memory replay method for continual learning in large language models.
Adaptive target reformulation approach reducing training instability in on-policy knowledge distillation for language models.
Geometric stability framework for measuring robustness of neural network representations.
Counterfactual explanation generation using fine-tuned LLMs for health intervention design and sensor data augmentation.
Fission-GRPO: Reinforcement learning method teaching LLMs to recover from tool execution errors in multi-turn interactions.
Cross-modal fine-tuning optimization technique balancing feature alignment and target fitting for pre-trained models.
Instruction purification method for improving reinforcement learning efficiency in LLM reasoning tasks.
Rate-distortion framework for lossy compression of transformer intermediate representations during inference.
Research on catastrophic forgetting in continual learning using mechanistic interpretability framework.
UniComp unified framework evaluating LLM compression via pruning, quantization, and distillation across performance, reliability, and efficiency dimensions.
Open-source discovery engine for photonic and hybrid quantum machine learning with optimized simulation and model evaluation framework.
Recursive transformer architecture with multi-resolution mechanisms for learning hierarchical dependencies via looped parameter sharing.
MASPO algorithm for LLM reasoning combining gradient utilization, probability mass, and signal reliability to replace rigid trust region mechanisms in RLVR.
Study measuring misalignment between LLM benchmark performance and real-world impact on downstream educational tasks across multiple models.
RL exploration method using temporal representations to learn environment models without extrinsic rewards or tracking novelty directly.
Query-aware GNN inference system guided by LLMs for efficient computation on large knowledge graphs with variable complexity queries.
Theoretical analysis of how differential privacy in neural networks affects fairness and adversarial robustness via feature-centric framework.
Countdown-Code benchmark environment for measuring reward hacking in LLMs with verifiable rewards, enabling clear distinction between task solving and reward manipulation.
Bilateral decoupled decay method for soft clipping in LLM reasoning with verifiable rewards, addressing gradient divergence in GRPO-style optimization.
LLM-based graph kernel approach for text-rich graphs that preserves raw textual information in message passing instead of compressing to embeddings.
In-context symbolic regression framework for Kolmogorov-Arnold Networks to extract interpretable analytical expressions from learned functions.
Test-time RL method for LLMs using selective-complementary pseudo-rewards to improve reasoning when consensus voting is unreliable on dispersed answer distributions.
Study on using LLMs to generate portable patient embeddings from clinical time series for cross-hospital model deployment.
RAVEN: generative pretraining foundation model for electronic health records via recurrence-aware next-visit event prediction.
Comparison of LLM agents vs classical HPO algorithms (CMA-ES, TPE) for hyperparameter optimization on autoresearch testbed.
ARCS: amortized analog circuit generation using graph VAE and flow-matching models, producing SPICE-simulatable designs in milliseconds.
Analysis of multi-token prediction in LLMs showing it promotes learning of structured world model representations.
Categorical mathematical framework for formalizing deep learning architectures, handling broadcasting and component composition.
Data deletion scheme for predicting model behavior after excluding training data subsets, relevant to interpretability and privacy.
DMax: efficient diffusion language model decoding with aggressive parallelism via progressive self-refinement from masks to tokens.
Method to recover dynamic programming structure from distributional reinforcement learning training dynamics.
THEIA: Modular neural architecture learning complete Kleene three-valued logic rules without symbolic inference, achieving 99% accuracy on all 39 K3 operators.
PAC-learning theoretical analysis of sample complexity differences between chain-of-thought and end-to-end autoregressive generation in language models.
GCA framework combining curated climate dataset with agentic LLM pipeline for GCC country climate decision support and geospatial tool integration.
Diffusion sequence models for meta-learning robot dynamics through in-context learning, comparing deterministic and generative approaches to system identification.
Token Importance analysis in on-policy knowledge distillation identifying which token positions provide most useful learning signals for student LLMs.
Contribution Weighted Group Relative Policy Optimization for training LLM-based search agents via reinforcement learning with improved credit assignment.
Best-Arm Identification framework for autonomous reasoning and planning with LLM agents, addressing systematic evaluation biases in search expansion.
Tabular foundation models for molecular property prediction using in-context learning without task-specific fine-tuning, improving drug discovery applications.
Knowledge-transfer network for multimodal sentiment analysis that reconstructs missing modalities during training and testing phases.
Research on logit suppression vulnerabilities in LLM safety alignment, introducing SSAG method to systematically identify and manipulate safety mechanisms in large language models.
Memory-aware techniques for efficient LLM training with automatic configuration in resource-constrained environments.