From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics
Multimodal LLM framework for annotating broadcast television content. Domain-specific application with limited general relevance.
Multimodal LLM framework for annotating broadcast television content. Domain-specific application with limited general relevance.
Study of LLM reasoning modes in multi-agent negotiation simulations. Examines behavior reproduction vs optimal solving in agent interactions.
LLM-based agent framework that generates proof-of-concept tests to validate bug detection reports. Combines agents with automated testing.
Self-evolving memory system for LLM-based code generation on private libraries using execution feedback. Improves code generation with enterprise context.
Semantic search system for clinical notes at health system scale using embeddings. LLM application but domain-specific healthcare focus.
Benchmark for LLM memory retrieval precision showing current evaluations mask severe precision failures through complete belief dumps.
Survey of mathematical reasoning in LLMs covering benchmarks, architectures, evaluation methods, and open challenges in the field.
Pass@K optimization research for code generation improving test-time compute allocation by coordinating diverse sampling instead of independent draws.
Research on scaling reinforcement learning from verifiable rewards for agentic LLMs using synthetic task augmentation instead of human curation.
Research showing LLMs fail to verify source quality during multi-source synthesis despite detecting fabrication in isolation.
TLA-Prover: 20B parameter LLM trained via preference optimization to generate formally verifiable TLA+ specifications for distributed systems.
MetaConfigurator extends JSON Schema editor with RDF authoring for semantic interoperability in scientific workflow data.
Training-free concept detection and steering in transformer models by analyzing sign patterns in raw transformer dimensions without learned dictionaries.
Study analyzing variability loss in AI-generated code from LLM vibe coding, showing programs have minimal compile/runtime variability.
Polycepta improves multi-object tracking with dynamic object-centric appearance estimation to complement motion prediction.
RWGBench introduces benchmark for evaluating related work generation in academic papers using citation-level scholarly positioning metrics.
Empirical study of ERC-8004 decentralized AI agent protocol, analyzing trust mechanisms in permissionless agent economies.
JuZhou 1.0 is ultra-lightweight text-to-image model designed for edge deployment and offline execution on China-developed AI accelerators.
MultAttnAttrib provides training-free multimodal attribution for long-document QA, improving interpretability and safety in AI assistants.
RoboDojo provides unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks.
Wan-Streamer v0.2 improves resolution of native-streaming audio-visual interaction model while maintaining 200ms latency.
SearchGen uses agentic visual generation with web search to overcome knowledge cutoff limitations in text-to-image models for open-ended user requests.
Weak-to-Strong Generalization enables cheaper RL post-training by running RL on small models then distilling to larger models for improved reasoning.
Floor-First Triage proposes analytical estimation methods for LLM serving optimization to replace grid-search approaches and improve latency profiling workflows.
UBEP optimizes Mixture-of-Experts model communication on high-bandwidth superpods by addressing execution serialization and bandwidth bottlenecks in production deployments.
TriRoute jointly optimizes attention resolution, expert selection, and KV-cache allocation in language models using learned routing to decouple model quality from per-token inference cost.
Open-source foundation model for wearable motion sensing with comprehensive study of pretraining and scaling principles.
LLM-guided time-series forecasting approach leveraging process documentation for industrial soft sensing with scarce labeled data.
Method for adapting specialist industrial models to new scenarios using LLM-guided reasoning without parameter modification.
Mechanistic study of reward valuation in vision-language models linked to anhedonia assessment from clinical psychology.
Approximation ratio analysis for greedy algorithm in myopic Bayesian active learning for linear regression with tight bounds.
DsrFGW: Optimal transport method for graph matching combining node features and structure via diffusion-inspired approach for sparse/noisy graphs.
Analysis of latent reasoning faithfulness in hidden state reasoning across training trajectories, showing unfaithful behaviors beyond converged checkpoints.
MESH-FL: Entropy-guided tensor compression for multimodal federated learning on edge devices accounting for modality-specific spectral differences.
Generative method for temporal point processes using rough path signatures as feature maps, addressing sequence-level evaluation limitations.
FedDualAtt: Personalized federated learning for ECG classification using split transformer attention heads with global and local branches.
Study on human-AI complementarity under asymmetric information, analyzing when human decision makers fail to realize gains from ML model augmentation.
Meta-learning approach for learned optimizers that efficiently scales to long-horizon inner problems, improving upon hand-designed optimizers like Adam.
Bayesian deep ensemble method for predictive regression combining statistical rigor with scalability and calibrated uncertainty estimates.
Knowledge distillation for time series classification, transferring knowledge from large teacher to efficient student model for resource-limited environments.
Analysis of signals predicting correctness in text-to-SQL generation using self-consistency and schema-relevance metrics on BIRD and Spider benchmarks.
Generative diffusion models for sampling stochastic signals on graphs, applied to recommender systems and financial forecasting.
LEMUR 2: Large-scale extensible NAS benchmark with 14,000+ distinct architectures and 750,000+ training records for cross-domain neural architecture evaluation.
Empirical study using rule-based expert as benchmark to evaluate lightweight RL agents for imperfect-information card games across 100+ experiments.
On-policy self-distillation method for LLMs using geometric approaches where teacher model has access to solution hints to improve student model reasoning.
Best-arm identification algorithm that pairs costly reward observations with cheap proxy scores from LLMs to improve data-driven decision-making efficiency.
Self-supervised image clustering framework using evolutionary algorithms instead of gradient descent, eliminating need for predefined loss targets.
DNN architecture learning for edge devices using zeroed batch normalization to meet strict latency constraints in real-time applications.
Wearable foundation model using physical activity data for scalable broad-spectrum health prediction.
Federated learning approach for rapid model adaptation under real-world client churn in recommendation systems.