SatSOM: Saturation Self-Organizing Maps for Continual Learning
Self-organizing maps extension addressing catastrophic forgetting in continual learning through saturation mechanisms.
Self-organizing maps extension addressing catastrophic forgetting in continual learning through saturation mechanisms.
Benchmark evaluating memory capabilities in LLM agents including memorization, updating, and retrieval of long-term information across multi-turn interactions.
Protocol-agnostic tool management library for function-calling LLMs that reduces fragmentation and development overhead in LLM applications.
Method for predicting better pre-trained weights via retrodiction of forgetting to encapsulate more knowledge beyond training datasets.
Analysis of fast weight programmers and 2D-state RNNs as linear transformers with connections to neurobiology and language modeling advances.
Framework for optimizing content for generative search engines powered by LLMs and RAG by modeling user intent and role-based search patterns.
Study showing psychometric questionnaires designed for humans mischaracterize LLM psychology compared to behavior in real user interactions.
Object-centric evaluation framework for automated fine-grained assessment of multi-turn instruction-based image editing using VLMs.
Tree-based group relative policy optimization for LLM agent reinforcement learning addressing sparse supervision in long-horizon multi-turn tasks.
Method aligning supervised fine-tuning with in-context learning activations to improve LLM generalization and calibration in data-scarce settings.
Offline RL framework using linear Transformers for compositional Q-function estimation across diverse subtasks via in-context learning.
Parameter-efficient machine unlearning method for foundation models addressing weight unbounding in privacy removal.
Benchmark evaluating metrics and judges for assessing harmful content generation in LLMs.
Slow-Fast Policy Optimization framework improving LLM reasoning via RL with better gradient stability and exploration.
Methods for detecting data contamination in LLM evaluation during reinforcement learning post-training phases.
Scalable energy-based models via adversarial training unifying classification and generative modeling using EBM framework.
CBF-RL integrates control barrier functions into reinforcement learning training for enforcing safety constraints during policy training.
Frame semantic patterns methodology for identifying underreported gender-based violence in e-medical records using NLP.
Generative hints training methodology enforcing functional invariances in vision models beyond empirical training data.
Neural network approach to accelerate computation of n-particle reduced density matrices for strongly-correlated quantum states.
Methodology for sustainable ML model evaluation addressing gaps in Green AI auditing practices and standardization.
Research showing genomic sequence models exhibit in-context learning similar to LLMs, demonstrating ICL emerges across sequence domains.
Uni-DAD: Unified method for distilling and adapting diffusion models for few-step image generation in new domains.
Benchmark for evaluating multimodal LLMs on schema-grounded visual information extraction and reasoning tasks in agentic settings.
Framework for runtime monitoring of multi-agent systems using cryptographic provenance and drift detection to mitigate emergent norms at scale.
Analysis of KL divergence estimators used in RL training of LLMs, evaluating different approximation methods for reverse KL regularization.
Benchmark for evaluating vision-language model routing systems across quality and cost dimensions using real inference logs.
Study on feature-dependent noise in preference-based reinforcement learning, examining how observation-dependent uncertainty affects learning.
Diagnostic benchmark for evaluating epidemiological reasoning in large language models, focusing on evidence-grounded inference over clinical knowledge.
Benchmark for evaluating frontier AI models on real-world software engineering tasks including integration and end-to-end system construction.
Reformulation of supervised fine-tuning to reconcile post-training objectives in large reasoning models using Gibbs initialization before reinforcement learning.
Systematic evaluation of speculative decoding acceleration techniques on production-grade vLLM inference engine, assessing real-world effectiveness.
Self-supervised reinforcement learning approach for improving low-resource machine translation using round-trip bootstrapping with NLLB models.
Analysis of YOLO26 object detection framework that eliminates NMS post-processing for native end-to-end learning strategy.
Method to improve vision-centric capabilities of multimodal LLMs through context-aware image representation prioritization via ensemble approach.
Case study examining why bias mitigation techniques fail on government data, using crime rate prediction with Bristol City Council data as test case.
Unsupervised method for decomposing complex data into factorized components using diffusion models, enabling reusable component recombination for images and robotic videos.
Knowledge distillation framework from vision-language models into lightweight networks for fine-grained visual classification using prompt-aware semantic calibration.
Graph domain adaptation method using neural characteristic functions to align distributional shifts across labeled source and unlabeled target graphs.
Framework for test-time verification of LLM reasoning traces using single-trajectory monitoring to catch errors without exploring multiple branches.
Zero-shot video anomaly detection framework using multimodal large language models without requiring real anomaly examples, addressing dataset diversity limitations.
Method for inferring hyperparameter trajectories using optimal transport to adapt neural network behavior post-deployment without expensive retraining.
Research on continual learning in pretrained Vision-Language-Action models for robot policy learning, showing resistance to catastrophic forgetting when acquiring new skills.
Benchmark evaluating LLM agent capabilities in long-term codebase maintenance via continuous integration, beyond static bug fixing.
Technique to reduce KV cache in transformers by using low-dimensional attention selection for keys while maintaining full-dimensional values.
LLM-guided GPU architecture exploration for LLM inference workloads via bottleneck analysis and design space optimization.
FRAME methodology for systematic evaluation of AI systems in real-world organizational contexts beyond abstract capability measures.
Method for achieving ethical fairness in AI systems without relying on demographic attributes in human-centered applications.
Study on software engineers' cognitive engagement with agentic coding assistants, examining over-reliance risks and critical thinking impacts.
Open-source biomedical knowledge graphs (pathways, clinical trials, drug interactions) with federated access and AI agent integration.