When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
Investigation showing 8-bit quantization paradoxically improves continual learning in LLMs by mitigating catastrophic forgetting compared to FP16.
Investigation showing 8-bit quantization paradoxically improves continual learning in LLMs by mitigating catastrophic forgetting compared to FP16.
NeuroFilter activation-based privacy guardrails for LLM agents controlling sensitive data access while maintaining agentic capabilities.
PaAno patch-based representation learning for time-series anomaly detection with lower computational cost than transformer baselines.
OmniMoE system-algorithm co-design for scaling Mixture-of-Experts with vector-level atomic experts improving parameter efficiency at scale.
Calibrated guidance mechanism for diffusion models addressing miscalibration in Bayesian posterior sampling using test-time guidance.
Token reduction technique for long-video vision-language models using Mamba-Transformer hybrid architectures with stateful compression.
Analysis of neural network visual cue preferences showing instability in stylization-based cue-conflict benchmarks for measuring shape bias.
Skeleton-based action recognition using Gaussian splatting and probabilistic topology for sensor-based human-computer interaction.
Method to improve LLM performance by maximizing mutual information between prompts and responses without additional training data.
Multi-Dependency PIBT: Enhanced multi-agent path finding algorithm for planning hundreds of agents in congested environments.
Knowdit: AI agent system for smart contract vulnerability detection using DeFi-specific semantic knowledge summarization.
EgoSim: Closed-loop egocentric world simulator generating interaction videos with persistent 3D scene state updates.
Presidio-hardened-x402: Middleware filtering PII from agentic payment metadata before transmission in x402 protocol.
System for generating scientific hypotheses from evolving literature using LLMs to identify promising research directions.
GUI-Perturbed: Framework revealing brittleness in GUI grounding models through domain randomization and controlled perturbations.
FED-FSTQ: Federated fine-tuning of LLMs on edge devices using Fisher-guided token quantization for communication efficiency.
SHIELD: Clinical NLP dataset and distilled small LMs for de-identification of health records at enterprise scale.
EcoGEO: Framework for understanding how LLM web-search agents are influenced by evidence across multi-step browsing and query trajectories.
ChainCaps: Framework for safe tool-using AI agents via monotonic capability attenuation to prevent unsafe multi-step tool compositions.
MIND: Diffusion model framework for image generation that explicitly models data manifold geometry with patch tokenization.
FaithRewriter: Prompt rewriting system for text-to-image generation that uses visual grounding to reduce intent-generation gap.
LC-QAT: Quantization-aware training method for 2-bit LLM compression using linear-constrained vector quantization to reduce model size.
Evaluation of generative AI for greenfield software engineering and 'vibe coding' practices without underlying domain knowledge.
Neuron-wise sequence modeling framework enabling independent neural evolution through topological dynamics instead of layer-wise constraints.
Evaluation of eight LLMs on mental health safety across DSM-5 conditions with adversarial attacks and harm taxonomy framework.
CAMS system for multi-document summarization with fine-grained claim-anchored attribution reducing hallucination in LLM outputs.
Evaluation of vision-language models for detecting synthetic medical images with text overlays and metadata in clinical contexts.
Certification mechanism for automatic speech recognition improving robustness to adversarial and benign perturbations without oracle knowledge.
Large-scale simulation of acoustic adversarial attacks against voice control systems scaling from digital to physical domain.
Framework using LLMs and vision-language models for spatial-temporal semantic search and recommendation of geographic information.
Large-scale bugfix benchmark using LLMs to generate diverse code corruptions reflecting real-world bug distributions.
Training-free gating method for few-shot reranking that uses model uncertainty to determine when reranking degrades LLM performance.
LLM routing system for regulated industries using classifier gates to enforce compliance and optimize inference cost on sensitive queries.
3D trajectory guidance for hierarchical vision-language action models in robot manipulation using depth-aware planning and control.
Benchmark for autonomous LLM-based financial agents measuring behavioral mandate decay over time under market context accumulation.
Representation framework for mechanistic interpretability enabling reusable, composable neural network component analysis and natural language querying.
Physics-constrained generative modeling using sparse nonlinear projection for inference-time constraint enforcement without retraining.
Framework disentangling classifier tuning and joint optimization for semi-supervised security classification pipelines.
Mixture-of-generators approach for synthesizing survival analysis training data in privacy-constrained clinical settings.
Analysis showing GRPO, Dr. GRPO, and DAPO LLM training methods adjust a single metric: standard deviation of model disagreement.
Neural architecture search framework using evolutionary algorithms to design task-adaptive Transformer models for time-series forecasting.
Parameter-efficient fine-tuning method using mixture-of-experts with learned transformation domains for model adaptation.
Reinforcement learning approach using verifiable scoring rules to train calibrated probabilistic forecasting models.
Scalable training algorithm for Ising-model-based thermodynamic computing devices for low-power AI inference.
Federated learning optimization reducing communication bandwidth for model parameter averaging and knowledge distillation across distributed peers.
Generates counterfactual feedback from superhuman game agents by analyzing latent geometry of expert performance.
Weak-form kernel ridge regression approach improves noise robustness for learning complex dynamical systems.
Benchmark for validating causal abstraction metrics across ten complex systems with ground-truth explanations.
Entropy-regularized probabilistic gates learn sparse models in federated learning under data heterogeneity.
Four-stage diagnostic evaluates LLM physics reasoning through induction, formulation, prediction, and review in unfamiliar frameworks.