Does a token buy you more or less now than it did a few months ago?
First-of-its-kind analysis of thousands of conversations measuring token value changes and invisible failures in coding assistants over recent months.
First-of-its-kind analysis of thousands of conversations measuring token value changes and invisible failures in coding assistants over recent months.
Research paper analyzing macroeconomic productivity effects of LLM coding tools across generations, showing gaps between token generation claims and actual shipping velocity.
Google upgrades NotebookLM with agentic reasoning capabilities, code execution, and multiformat output generation for research projects.
Best practices guide on using Claude effectively by rejecting auto-agreement, using plan mode critically, and ensuring conceptual soundness before execution.
Disentangled rectified flow approach for time series super-resolution to reconstruct high-resolution signals from low-resolution inputs.
LEAF reinforcement learning method for speech-aware LLM post-training using tree-based token credit assignment instead of coarse rewards.
Multi-armed bandit algorithm for structured neuron pruning in deep neural networks to improve efficiency and parameter reduction.
Item Response Theory framework for estimating LLM scaling laws efficiently without expensive checkpoint evaluations across thousands of models.
Query Lens extends Logit Lens to interpret sparse autoencoder features by analyzing encoder and decoder-side key-value pairs to understand feature activation and output promotion.
ScaleSweep improves NVFP4 4-bit quantization of LLMs through efficient block scale optimization, reducing gap to optimal solutions.
HASA allocates subnets in federated learning to handle heterogeneous device resources and data distributions while optimizing statistical performance.
Model-theoretic framework with finite semantic certificates for verifying context-conditioned LLM behavior and understanding emergence through row-space criteria.
Sequential statistical inference framework for LLM trustworthiness, modeling dependent stochastic processes in deployment with behavioral monitoring.
Position paper arguing LLMs should optimize for personalized individual preferences rather than aggregated preferences that represent no real user well.
Active learning framework using foundation model priors to mitigate class imbalance and reduce annotation costs on imbalanced datasets.
Trait-space monitoring detects emergent misalignment during LLM supervised finetuning using linear directions in activation space without behavioral evaluation.
Framework for evaluating ML resource utilization and environmental impact across full model life cycle from development through deployment.
KITE integrates text, images, and knowledge graphs in a tri-modal transformer for fake news detection with improved semantic consistency analysis.
DOG-DPO optimizes safety alignment for LLMs through dynamic data selection that preserves directional preference information across multi-dataset settings.
Semantic Cache Distillation optimizes LLM inference by reducing KV cache communication overhead and enabling cache reuse across model variants.
Test-time adaptive composition framework for ML-as-a-service in IoT environments handling heterogeneous client resources and data distributions.
HARP presents efficient data selection for LLM finetuning, balancing scalability and downstream utility through train-free and train-based selection methods.
First systematic evaluation of activation steering robustness in LLMs under adversarial text perturbations, testing four extraction methods and three attack strategies.
Proposes surrogate-assisted evolutionary algorithm for client selection in federated learning to improve convergence and robustness.
Research on optimizing dense attention in long-context LLMs using oracle-guided sparse prefill to reduce computational costs while preserving task performance.
FunctionEvolve method combining LLM guidance with symbolic structure for symbolic regression, discovering scientific laws from data with domain-informed search.
SAW stage-aware dynamic weighting technique for multi-objective RL in LLMs, addressing asynchronous reward learning across objectives.
WhiFlash accelerates speculative decoding via token-level cross-paradigm routing, switching between autoregressive and diffusion drafting based on token type.
Rosetta Memory system providing adaptive memory for multi-LLM agents, enabling persistence and experience accumulation across different LLM backends.
SySRs bandit algorithm reducing LLM evaluation costs by exploiting model similarity and adaptively allocating evaluation budget.
Study showing contrastive learning methods confuse slow noise with dynamics, proposing fixes for predictive representation learning frameworks like JEPA.
Framework and benchmarking suite for evaluating concept drift detection methods in data stream mining with standardized metrics.
Study of Byzantine adversarial resilience in multi-agent LLM coordination games, examining robustness of communication protocols under attacks.
Teacher-free self-training study examining whether language models acquire new capabilities or improve expression of existing ones using exact verifiers.
Instrumented data approach for scientific machine learning combining observational data with mechanistic models and uncertainty quantification.
STILL method for KV cache compaction in language models during inference, addressing memory bottleneck for long-horizon deployment.
Pipeline parallelism optimization technique (PACI) for training large neural networks with bounded weight inconsistency, reducing memory bubbles.
Research on strained coherence failure mode in LLM-based coding agents where agents acknowledge problems but proceed anyway, related to reward hacking.
Framework studying feedback loops when ML models deployed in real-world settings, examining endogenous distribution shifts.
Contextual bandits approach for selecting active learning strategies dynamically without relying solely on labeled data feedback.
RL method for improving LLM reasoning via verifiable rewards, analyzing confidence inflation and inefficient compute allocation in standard GRPO training.
Cross-domain minibatch selection method using gradient matching and partition matroid constraints for training LLMs on heterogeneous data.
Physics-informed neural network framework for solving high-dimensional heat diffusion under noisy boundary conditions.
Discussion of localized architectures for improving interpretability and safety of large language models and reasoning models.
Study investigating task ordering effects on catastrophic forgetting in continual learning, comparing general-to-specific versus mixed learning orders.
Semantic reliability certification framework for multi-agent LLM systems in autonomous cloud operations, addressing operationally unsafe mutations.
Privacy-preserving vertical federated learning using causal representation learning to defend against sample reconstruction attacks.
Semi-supervised learning approach for ECG classification handling label scarcity and out-of-distribution anomalies in clinical settings.
Analysis of LLM safety evaluation gaps, formalizing audit gap between behavioral safety and representation-level robustness under intervention.
Study of graph reconstruction attacks on GNNs and defenses against model inversion attacks that reconstruct training adjacency from trained models.