Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology
Geometry-aware uncertainty coresets for robust few-shot in-context learning with vision-language models on histopathology images.
Geometry-aware uncertainty coresets for robust few-shot in-context learning with vision-language models on histopathology images.
Comprehensive survey of mathematical reasoning in LLMs covering benchmarks, architectures, evaluation methods, and open challenges.
MambaGaze framework using bidirectional Mamba for cognitive load assessment from eye-tracking data with missing data handling.
MechVQA benchmark for evaluating multimodal LLMs on mechanical engineering drawing understanding and spatial reasoning tasks.
LLM-native psychometric instrument reveals gap between model self-reports on personality dimensions and actual behavioral patterns across 25 models.
SEVRA-BENCH benchmark testing vulnerability detection in LLM-based code review agents against social engineering attacks on pull requests.
Agentic judging pipeline using LLMs for scalable architectural evaluation of code, improving code LLM quality beyond functional correctness.
LLM-driven information extraction system for tracking illegal fishing, seafood fraud, and labor abuses in supply chains.
Study of LLM agent tool selection behavior revealing over-privileged tool escalation despite lower-privilege alternatives being sufficient.
Neuromorphic reinforcement learning framework for pathfinding optimization in robotic warehouse systems with real-time constraints.
AutoSpec framework for evolving safety rules in LLM agents via inductive logic programming, balancing interpretability with operational flexibility.
Research showing task-conditioned language/vision models suppress reporting of safety-critical signals present in data, analogous to human inattentional blindness.
LLMs applied to threat extraction from media for peacekeeping mission risk assessment using OSINT collection and structured information mapping.
Efficient linear attention mechanism with content-aware recurrent updates and value efficiency for chunk-parallel processing.
Open-source Mathswitch project uses LLM voting ensembles to categorize and link mathematical concepts from multiple sources.
Ember optimizer exploits embedding table geometry to improve LLM fine-tuning, RL, and pretraining with minimal optimizer state.
Study of how safety alignment in LLMs creates overly restrictive constraints in cybersecurity domains, examining domain-specific refusal patterns.
Fine-grained entropy-based method for improving token-level credit assignment in RL-trained LLMs, addressing sparse reward challenges.
Multi-agent reinforcement learning framework for energy management in AI data centers with carbon-awareness optimization.
Study of transferability between image understanding and generation tasks in unified multimodal model architectures.
RoboDojo: unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks.
Wan-Streamer v0.2: upgraded streaming audio-visual interaction model achieving 640x368 resolution at 200ms latency.
TACTIC-KG: multi-agent LLM system constructing cyber threat intelligence knowledge graphs from unstructured CTI reports.
Audex: unified audio-text LLM built on Nemotron MoE encoding audio into text embedding space without degrading text performance.
First multiplayer world model for dynamic environments conditioning on multiple agents' action streams with scene coherence.
Design-CP: context parallel inference strategies for RFdiffusion 3 protein design using row/grid sharding with ring attention.
Exogenous dropout technique for robust time series forecasting with noisy or missing exogenous covariates.
DNN compression via controllability-observability framework for state-order reduction treating networks as dynamical systems.
Method for learning to control LLM agent execution harnesses via offline reinforcement learning while keeping LLM frozen.
Study of parameter-free encoders for relational database foundation models to predict missing values across varied prediction tasks.
InvWeaver: neuro-symbolic framework using LLMs for loop invariant synthesis in multi-loop programs via deductive feedback.
PatchOptic system for LLM agentic workflows using projected views and verified structured updates over shared state.
Self-review RL with cross-episode memory for training LLMs with sparse/delayed feedback and policy distillation.
FedPPO-PG combines federated learning with physics-grounded PPO for multi-agent stability control in smart grids.
Conformal prediction method for reliable clinical data imputation with uncertainty quantification for high-stakes decisions.
Stochastic sparse token steering for LLMs using probabilistic gating to reduce per-token perturbation overhead.
Safe Bayesian optimization with counterfactual policies ensuring interventions don't degrade outcomes below baselines.
Deep reinforcement learning for autonomous mobile robot battery charging optimization in warehouse environments.
LLM-guided generation of neural network architecture improvements using same-family source models as guidance.
FourTune: 4-bit quantization method for efficient post-training fine-tuning of diffusion models with low memory overhead.
Training-time defense against neural backdoor attacks by learning to distinguish poisoned samples from benign ones.
In-context learning applied to antibody affinity ranking for drug discovery by capturing antigen-specific binding landscapes.
arXiv paper on multi-buyer negotiation using reinforcement learning with LLMs for strategic bargaining with private information.
arXiv paper on differentially private natural gradient descent improving optimization efficiency under privacy constraints.
arXiv paper demonstrating non-identifiability of subspaces in low-rank LLM training, challenging GaLore optimizer assumptions.
arXiv paper proposing practical auditing of unlearning algorithms using membership inference attacks to verify data removal.
arXiv paper introducing K-ABENA, a selective gradient computation framework reducing training costs via sample exclusion and unbiased estimation.
Research paper showing LLM judges optimize for plausibility over correctness in self-play settings, enabling reward hacking.
arXiv empirical study of neural architecture robustness to temporal distribution shift across time-indexed domains.
arXiv paper on recovering sparse linear causal DAGs with latent confounders using higher-order cumulants.