kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail
kNNGuard: Training-free guardrail for LLMs using activation space and k-NN to detect unsafe/adversarial prompts without fine-tuning.
kNNGuard: Training-free guardrail for LLMs using activation space and k-NN to detect unsafe/adversarial prompts without fine-tuning.
Bayesian active ranking method using LLM judges to identify top-k candidates while accounting for systematic biases and position effects.
ART: continuous-time control framework using actor-critic learning to optimize timestep allocation in score-based diffusion sampling.
DALorRA: Bayesian sparse low-rank adaptation method for uncertainty quantification in fine-tuned LLMs to improve trustworthy deployment.
Privacy-preserving distributed coded computing framework addressing privacy leakage and malicious manipulation in federated and decentralized learning.
HERMES provides hierarchical multi-granularity labeling system for organizing pre-training data mixtures across different semantic axes.
Study of generalization in offline reinforcement learning showing structure of pessimism matters more than degree for contextual MDPs.
Framework for training visual generative models using distribution-wise rewards to prevent reward hacking and improve image diversity.
DecompRL uses reinforcement learning to teach LLMs modular code generation for solving hard problems by decomposing into solvable subcomponents.
Active few-shot learning method for LLMs that identifies valuable unlabeled samples for annotation to reduce human labeling costs and improve domain-specific adaptation.
Federated learning with quantum enhancement for multi-agent activity recognition in distributed robotic systems addressing non-IID heterogeneous sensor data.
Theoretical analysis of distributed self-supervised learning robustness under non-IID data heterogeneity in decentralized settings.
Neuron-aware data selection method for annotation-free LLM self-distillation in specialized domains without human-labeled supervision.
Compares alternative optimizers (SOAP, Muon) to Adam for training machine learning interatomic potentials for scientific simulation.
On-policy self-distillation for LLMs using disagreement-modulated approach to improve reasoning while reducing overfitting and improving cross-domain generalization.
Fuzzy-function programming paradigm compiling natural language specifications into locally-executable neural artifacts as alternative to LLM APIs.
Office Comprehension Bench: first benchmark for evaluating LLM systems on Word, Excel, and PowerPoint document understanding.
Theoretical analysis of ReLU neural network approximation for binary classification over o-minimal definable sets.
Benchmark comparing 13 federated learning and 10 knowledge distillation algorithms for 3D point cloud classification on edge devices.
X-VAE: variational autoencoder framework learning data-adaptive Gaussian mixture priors instead of standard isotropic priors.
Research on black-box embedding inference attacks against dense IR systems without knowledge of target embedding models.
Research on few-shot audio classification handling unseen classes with attention-based prototype methods.
Survey of generative AI and federated learning approaches for intrusion detection systems in IoT and distributed networks.
Research paper on quantum cost landscape geometry and optimization paths in variational quantum algorithms using nudged elastic band methods.
Research paper analyzing neural quantum states using sparse autoencoders for mechanistic interpretability and causal feature steering.
Bi-NAS: bi-level neural architecture search for generating personalized and effective explanations in recommender systems.
GPUAlert: zero-instrumentation process-boundary monitor for diagnosing GPU training job failures without modifying training scripts.
BIFROST: sim-to-real transfer method for robot policy learning that learns invariant feature representations addressing both visual and kinematic domain gaps.
CreativityNeuro: data-free method using contrastive weight steering to enhance divergent thinking in LLMs and reduce mode collapse on open-ended generation.
Adapts mixture-of-experts diffusion language model DiffusionGemma-26B for medical radiology report generation and benchmarks against autoregressive baseline.
Procedural Memory Distillation: method for language models to retain and reuse procedural information across episodes for self-improvement through online reflection.
Semi-CoT: framework for semi-supervised chain-of-thought learning that reuses generated reasoning traces as learning signals to improve LLM reasoning capabilities.
Monograph introducing mean field reinforcement learning through Markov decision processes and large-population stochastic control with mathematical framework.
OPINE-World: programmatic world modeling using LLMs and counterexample-guided synthesis to generate data-efficient, reusable environment models for agent adaptation.
Proposes MMAO-Cls using metabolic multi-agent optimization as outer-loop optimizer for joint feature selection and classifier hyperparameter tuning.
AgenticDataBench: benchmark for evaluating LLM-based data agents on automating data science workflows including data wrangling, analysis, and visualization tasks.
Studies adversarial robustness and explainability stability of cybersecurity classifiers using SHAP-based explanations across multiple datasets and attack methods.
Introduces Goggles, a learned module using gradient editing to improve language models' ability to recognize fictional content, addressing the negation neglect problem.
COMFYCLAW: agentic system with self-evolving skill harnesses for image generation workflows, enabling agents to recall patterns and user preferences from prior runs.
Full Bayesian reinforcement learning approach via Likelihood-Free Iterative Bayesian Importance Sampling for data-scarce settings.
PARTREP method enabling decoder-only LLMs to learn selective prompt repetition patterns, improving reasoning by redistributing contextual grounding across positions.
Lynx: progressive speculative KV cache quantization technique for accelerating long-context LLM inference in retrieval-augmented generation and agentic systems.
PhysMani framework coupling physics-principled 3D Gaussian world model with action policy for dynamic object manipulation in embodied AI.
Object Aligner: configurable JSON schema similarity scoring for measuring LLM output alignment with structured schemas, enabling agentic planning and tool calling evaluation.
Evaluation of Vision-Language Model reliability for medical image quality assessment under image corruption and demographic bias.
Maven RL framework with editable evidence memory for long-context reasoning, rewarding intermediate evidence state changes rather than just final answers.
Open-weights constitutional classifier for multilingual AI safety filtering, achieving SOTA on prompt-safety benchmarks at 1/10th the size of competing models.
Emotional Self-Correction method improves vision-language model reliability by activating latent self-correction without post-training.
Evaluates frontier LLMs on expert-authored clinical reasoning scenarios, showing open-ended medical performance remains unsolved with 32% hard subset score.
Proposes Fourier-based preconditioning for mutual information-inspired feature learning, proving H-Score invariance properties.