Algebraic Quantum Intelligence: A New Framework for Reproducible Machine Creativity
Algebraic quantum intelligence framework proposing quantum computing approach to improve creative output generation in LLMs.
Algebraic quantum intelligence framework proposing quantum computing approach to improve creative output generation in LLMs.
DenseMLLM: framework enabling multimodal LLMs for dense prediction tasks like segmentation and depth estimation without task-specific decoders.
Fast image and video editing using diffusion model guidance without costly vector-Jacobian product computations.
CORPGEN: multi-horizon task environment benchmark for evaluating autonomous agents on concurrent long-horizon tasks with dependencies and reprioritization.
Sali-Cache: dual-signal KV-cache optimization framework for efficient long-form video understanding in vision-language models.
Federated ensemble learning approach with progressive personalization addressing statistical heterogeneity across distributed clients.
GRAIL: imitation learning approach for goal recognition alignment enabling accurate identification of agent goals from behavior for AI alignment.
AD-Bench: real-world benchmark for evaluating LLM agents on complex advertising and marketing analytics tasks requiring multi-round reasoning.
STATe-of-Thoughts: interpretable inference-time compute method using structured action templates for improved diversity and explainability in LLM reasoning.
SM-EM algorithm for ML optimization reformulating EM iterations as weighted least squares with learnable scaling analogous to Adam optimizer components.
FMMD: Multimodal peer review dataset from F1000Research for training AI systems for automated scholarly paper evaluation.
Floe: Federated learning framework combining cloud LLMs with edge small language models for low-latency, privacy-preserving real-time inference.
Analysis of LLM benchmark saturation problem: frontier models exhaust new benchmarks quickly, threatening ability to measure AI progress.
Coalition formation model for selfish agents with possibly overlapping coalitions under partial information using offline learning.
AdaptManip: Reinforcement learning framework for humanoid robots to autonomously perform navigation, object lifting, and delivery without demonstrations.
InnoEval: Framework for evaluating AI research ideas using LLMs with knowledge-grounded, multi-perspective reasoning and collective deliberation.
LRD-MPC uses low-rank decomposition to improve efficiency of secure multi-party computation for machine learning inference.
PITA dataset with 23M propositional logic statements examining how reasoning traces support LLM reasoning and length generalization limits.
Frontier AI Risk Management Framework analyzing risks from rapidly advancing AI models and agentic AI systems, version 1.5 technical report.
SWA: Game-theoretic framework for multi-agent LLM systems balancing individual alignment with collective stability through modified inference-time decisions.
Two-stage deep reinforcement learning approach for training quadruped robots to climb U-shaped stairs autonomously.
COOL-MC framework for formally verifying and explaining reinforcement learning policies for sepsis treatment using model checking.
Evaluates mathematical reasoning capabilities of LLMs in Sinhala and Tamil languages versus English-like translation representations.
TWISTED-RL framework for robotic knot-tying using hierarchical reinforcement learning agents without human demonstrations.
MATEO: multimodal benchmark for evaluating LVLMs on temporal reasoning and planning with directed acyclic graph task execution orders.
LongAudio-RAG: hybrid framework combining audio-language models with retrieval-augmented generation for multi-hour audio question answering.
VariViT: Vision Transformer architecture supporting variable image sizes without fixed-size patches, addressing medical imaging challenges.
Tabular foundation models applied to association rule mining, outperforming classical and neural approaches especially in low-data regimes.
Analysis of prefill attack vulnerability in open-weight LLMs, exposing systematic security risks in models relying on internal safeguards.
Evolutionary System Prompt Learning (E-SPL) method for jointly improving LLM contexts and weights through reinforcement learning iterations.
LLMStructBench: benchmark for evaluating LLMs on structured data extraction and JSON generation from natural language across 22 models and 5 prompting strategies.
Research on behavioral self-awareness in LLMs fine-tuned with incorrect data, examining how misalignment emerges and shifts with realignment.
Dataset paper on knowledge graph refinement algorithms with schema-level information and neurosymbolic techniques for ontological reasoning.
Study on pre-trained protein sequence embeddings for machine-learning-based protein design, addressing challenges with sparse mutation datasets in bioengineering.
RF-GPT extends LLMs and multimodal models to natively support radio-frequency signals for wireless systems, bridging gap between LLM-based telecom approaches and RF signal processing.
Multi-dimensional persistent sheaf Laplacian framework for image analysis addressing dimensionality reduction sensitivity.
Theoretical analysis of temperature scaling properties for classifier calibration and LLM stochasticity control.
Drift-diffusion matching framework embedding asymmetric RNN dynamics in latent manifolds for biological neural computation.
GAPA: post-hoc Gaussian Process method for uncertainty quantification in pretrained networks via activation-space modeling.
ThermEval benchmark for evaluating vision-language model generalization on thermal imagery across surveillance and medical applications.
Quantum-enhanced Gaussian processes with distributed training for multi-agent systems using quantum kernel embeddings.
BPP method for robot imitation learning attending to relevant historical frames via key frame selection in long-context tasks.
Cold-start personalization using structured world models and RL for efficient user preference inference with limited interaction budget.
Regression algorithm using subset stacking for heterogeneous data with local predictors trained on random input-space subsets.
Analysis of sample efficiency in offline policy selection for reinforcement learning via off-policy evaluation and Bellman error.
PIVID method for inferring Bayesian network structure from observational data using permutation-based variational inference.
Memory-efficient zeroth-order optimizer (MeZO) for LLM fine-tuning using only forward passes, with sparse parameter improvements.
Optimal design methods for efficient human preference elicitation to reduce annotation costs in preference learning models.
Survey of federated learning combined with foundation models for privacy-preserving distributed training.
Multi-graph pretraining approach enabling graph transformers to learn transferable representations across diverse domains.