GSVD for Geometry-Grounded Dataset Comparison: An Alignment Angle Is All You Need
GSVD method for comparing datasets using linear relations and co-span constraints in shared ambient space. Mathematical technique for geometry-grounded learning.
GSVD method for comparing datasets using linear relations and co-span constraints in shared ambient space. Mathematical technique for geometry-grounded learning.
GaLoRA: Parameter-efficient adaptation of LLMs for graph node classification on text-attributed graphs.
Regime-aware in-context learning framework using LLMs for financial volatility forecasting without fine-tuning.
Search procedure to identify optimal neural network learning rate schedule shapes for different training workloads.
Sampling strategy for masked language models optimizing protein properties, evaluated in silico and on antibody therapeutics.
Federated active learning method addressing extreme non-IID data and global class imbalance with query-model selection.
Causal Concept Graphs for interpretability in LLMs, combining sparse autoencoders with structure learning for multi-step reasoning.
Neural scaling laws for MoE architectures determining optimal compute allocation between expert and attention layers.
Variance-aware adaptive weighting strategy to balance diffusion model training across different noise levels.
Graph-GRPO: RL-based training method for graph flow models aligned with task-specific objectives for drug discovery.
Framework using LLMs to design service systems by analyzing textual evidence from customer interactions and operational reports.
Analysis of mean bias effects in FP4 quantization during LLM training, addressing numerical instability from anisotropic weight distributions.
Method for instance-level unlearning in diffusion models without text prompts, targeting specific undesired outputs.
Multi-agent RL framework for coordinating UAV fleets for time-critical medical supply delivery with resource constraints.
Addresses length inflation in LLM training via group relative reward rescaling in RL-based fine-tuning without performance trade-offs.
SCORE: alternative to layer stacking using recurrent ODE-inspired blocks for deep neural networks with skip connections.
RL approach to optimize cluster job scheduling by learning to weight multiple scoring functions for node allocation.
Investigation of in-context learning in transformers through statistical hypothesis testing, revealing likelihood-ratio test approximation mechanisms.
Hardware-aware post-hoc ensembling method balancing accuracy and computational efficiency for tabular data via Pareto optimization.
Reinforcement learning with conditional expectation rewards extending RLVR to free-form reasoning domains beyond verifiable reward scenarios.
Methods for reliability and uncertainty estimation in CNNs addressing poor calibration and overconfident predictions in deep networks.
Grammar-based structural framework decomposing supervised learning into 7 primitives with DAG constraints to prevent data leakage in ML workflows.
CUPID plug-in framework for joint aleatoric and epistemic uncertainty estimation in deep neural networks via single model.
Value-driven memory approach for LLM-based kernel synthesis on data-scarce NPU architectures, addressing cold-start problems without fine-tuning.
Generalist value models as priors for sparse RL rollouts in reinforcement learning with verifiable rewards, improving policy gradient training.
Novel online prompt selection method for RL finetuning of LLMs, focusing on dynamics-predictive sampling to improve reasoning abilities through better training data selection.
Analyzes ergodicity in RL optimization, showing expected value across trajectories differs from single long-trajectory performance, affecting deployment-stage agent behavior.
KV cache eviction technique for efficient LLM inference on long-context tasks using future token prediction.
Safety approach for RLHF using stochastic dominance for robust risk control beyond expected cost constraints.
Scorio library implementing statistical ranking methods for comparing reasoning LLMs under test-time scaling.
CVE-based benchmark evaluating LLM capabilities and vulnerabilities in software security tasks and development.
Analysis of transformer MLP layers revealing binary routing of continuous signals with specific neuron implementations.
Leech Lattice Vector Quantization: novel LLM compression technique using structured lattice approaches for efficient parameter encoding without explicit codebook storage.
Proposes ConFu, adaptive method improving speculative decoding for LLM inference acceleration by enhancing draft model token proposal quality.
Prototype RAG-based AI assistant for navigating internal documentation in large physics collaborations like CMS at CERN.
Applies parameter-efficient fine-tuning to LLMs for multitask code analysis, unifying diverse code-analysis objectives in single model.
Uses evolutionary optimization with chain-of-thought to discover effective feature transformations for improving downstream ML model performance.
Framework for building domain-adapted LLM conversational systems via supervised fine-tuning, RAG, and evaluation for institutional deployment.
Evaluates offline LLM robustness and pedagogical safety for Turkish heritage language education in privacy-constrained educational settings.
Theoretical investigation of emergent LLM properties including semantic comprehension, in-context learning, and chain-of-thought reasoning mechanisms.
Proposes using Wikidata to create sociocultural bias datasets for detecting LLM prejudices against Latin American cultures in non-English languages.
Introduces SpreadsheetArena platform for evaluating LLM performance on end-to-end spreadsheet generation from natural language constraints via pairwise evaluations.
Investigates whether LLMs can deceive without lying, challenging mechanistic approaches that rely on truth probes to detect deception in model outputs.
Fine-tuned multilingual E5 encoder for detecting AI-generated Arabic text, comparing pooling strategies with mean pooling achieving F1 of 0.75.
Hierarchical tiered memory system for long-running autonomous agents with importance-aware eviction and hybrid routing.
Transformer-based framework for fault detection and localization in heterogeneous IoT sensor networks for smart homes.
Evaluation of autonomous cyber attack agent generalization across unseen network IP configurations using meta-learning.
Large-scale study measuring how agentic scaffolds and deployment configurations affect measured LLM safety in practice.
Framework enhancing vision-language-action robot policies with flexible guidance sources for complex manipulation tasks.
Red-teaming study on adversarial semantic layer activation steering attacks and defenses for LLM safety.