Photons = Tokens: The Physics of AI and the Economics of Knowledge
Quantitative analysis of AI computation costs using thermodynamic principles, defining tokens as measurable physical quantities.
Quantitative analysis of AI computation costs using thermodynamic principles, defining tokens as measurable physical quantities.
SmartBench evaluates LLMs on anomalous device detection and reasoning in smart home environments beyond basic command execution.
HEARTS benchmark evaluates LLM reasoning on diverse health time series with complex temporal dependencies across multiple physiological modalities.
RECAP applies Hebbian learning and local plasticity to reservoir computing for image recognition without backpropagation.
SR-TTT improves test-time training language models by using surprisal-aware residual learning to fix catastrophic failures on exact-recall tasks.
Federated learning framework for medical imaging using trust-aware mechanisms to handle unreliable participants in distributed environments.
Compares LLMs and SLMs for Intent-Based Networking orchestration in 5G/6G using hierarchical multi-agent architecture.
ObjChangeVR uses multimodal LLMs to reason about object state changes from continuous egocentric VR views with sparse motion cues.
Geodesic gradient descent optimizer for learning on objective function-induced manifolds without manual learning rate tuning.
PaLMR framework aligns multimodal LLM reasoning at both process and outcome levels using reinforcement learning to reduce process hallucinations.
FCBNet, parameter-efficient CNN for weed segmentation in multispectral aerial imagery using frozen ConvNeXt backbone and feature correction blocks.
GameVerse benchmark evaluates vision-language models using reflect-and-retry paradigm in video game environments to assess learning from failure.
Graph-based visual prompting technique for multimodal language models that improves spatial reasoning by representing marked objects as graph structures.
Optimizes Diffusion Transformer video generation inference using sequential-parallel 3D positional encoding to reduce memory and latency bottlenecks.
Analysis showing chain-of-thought prompting underperforms in medical vision tasks due to perception bottlenecks in VLMs.
Graph-based self-imitation learning for orchestrating edge AI and microservices with latency constraints.
Parallel Relative Policy Optimization for improving chart understanding and deep research capabilities in large vision-language models.
Eye-tracking gaze trajectories used as supervision to improve vision-language model reasoning in medical radiology tasks.
Method for mining temporal specifications and data transformations from execution traces using syntax-guided synthesis.
ATLAS framework for reinforcement finetuning enables small language models to operate effectively over large tool ecosystems with long-horizon planning and sparse rewards.
ProtAlign: Contrastive learning framework aligning protein sequence and structure embeddings for improved protein language models.
Hybrid approach combining time series foundation models with regression for electricity price forecasting capturing temporal and cross-variate patterns.
Safe Transformer: Modular safety alignment via explicit discrete safety bit bottleneck for interpretable and controllable LLM behavior.
Reinforcement learning for safe navigation through variable-density crowds generalizing beyond training densities without freezing.
Agent Hunt: Multi-agent autoformalization system using bounty-based marketplace where LLM agents propose and prove lemmas in theorem proving.
Super-Resolution Transformer optimization removing relative positional bias constraints to enable FlashAttention and efficient inference.
ResearchEnvBench: Benchmark evaluating autonomous agents on environment synthesis and complex software dependency resolution for research code execution.
ViroGym: Benchmark for evaluating protein language models on viral protein variant effects prediction with in silico and in vitro validation.
Heterogeneous decentralized diffusion training framework reducing computational requirements while supporting diverse training objectives across distributed experts.
Framework for constrained generation bridging pretrained models to respect complex physical constraints in robotic control and autonomous driving.
Method stabilizing GRPO reinforcement learning for diffusion language models by addressing reward collapse and incompatibilities with dLLMs.
DIRECTER: Dynamic rejection steering method to enhance LLM instruction following while avoiding oversteering degradation via activation steering.
ButterflyViT: Expert compression technique achieving 354× compression for sparse Mixture of Experts Vision Transformers on edge devices.
Property-driven protein inverse folding with multi-objective preference alignment for sequence design balancing designability and developability.
Survey of robotic foundation models for industrial control, analyzing applicability, requirements, and use cases in collaborative robotics.
Symbolic machine learning approach for failure detection in chemical processes with interpretability focus.
Deep reinforcement learning with heterogeneous graph transformers for job shop scheduling.
Python package providing reusable infrastructure for evaluating attribution methods on synthetic time series data.
Mechanism for deep reinforcement learning that preserves exploratory behavior via historical trajectory buffering.
Research on mechanistic understanding of LLM misalignment behaviors including deception and manipulation.
Study on how temporal visual grounding in vision-language models affects out-of-distribution generalization.
Tool for discovering term abstractions in automated reasoning using pattern mining from proofs.
Mechanistic interpretability study identifying audio-specialist attention heads in audio-language models to improve audio grounding over text dominance.
C3 method for credit assignment in multi-agent RL systems with LLMs using counterfactual reasoning to improve learning from sparse terminal feedback.
Study using LLMs to automate artifact evaluation for published security research papers, improving reproducibility checking scalability.
Dual-Transformer-Cascade framework for trajectory prediction integrating environmental priors for flying objects with hardware efficiency considerations.
Physics-informed surrogate model for ferroelectric vertical NAND retention analysis, accelerating simulations from day-scale TCAD to second-scale predictions.
Multi-agent LLM system personalizing mindfulness meditation experiences through expert-aligned design to improve user engagement.
Privacy analysis demonstrating inversion attacks on DNA embeddings from foundation models shared via Embeddings-as-a-Service frameworks.
Theoretical analysis of policy gradient post-training for linear autoregressive models with outcome and process rewards on margin-separated sequences.