D$^3$-Subsidy: Online and Sequential Driver Subsidy Decision-Making for Large-Scale Ride-Hailing Market
Online sequential optimization for driver subsidies in ride-hailing using reinforcement learning.
Online sequential optimization for driver subsidies in ride-hailing using reinforcement learning.
Lottery-based selection mechanism for stable randomized decision-making in competitive processes.
Aligns GRPO reinforcement learning with state-action modeling for vision-language model agents in open-world tasks.
Method for achieving neural collapse in supervised classification through prototype learning on hypersphere.
Latent diffusion model for airfoil generation with geometric validity constraints and physical controllability.
Adaptive weighting method for random forests using decision-path patterns to improve classification.
Federated learning approach with proactive client selection to handle non-IID data and improve convergence.
Addresses policy divergence in RL by formalizing behavior-consistent training for reliable deployment.
Novel approach combining generative policies with entropic mirror descent for off-policy reinforcement learning.
Study of adversarial robustness in spiking neural networks with analysis of attack transferability.
Research on accelerated sampling from diffusion models using Gaussian Mixture Models for moment matching in DDIM.
Causal attribution model enhances LLM interpretability and reasoning through do-operators and precision fine-tuning.
ImProver agent automatically optimizes formal proofs in Lean by applying style, readability, and modularity criteria.
LLM-based quantification of internal narratives to characterize affective states in psychological assessment.
NaviAgent uses bilevel graph-based planning for tool orchestration in LLM agents, scaling to hundreds of tools with dependency handling.
Agentic planning framework using world models for simulative reasoning, enabling task generalization without re-engineering.
Study measuring efficiency of small language models on local hardware, comparing energy consumption and performance to cloud-based alternatives.
WarmServe improves multi-LLM serving on shared GPUs by predicting workloads and prewarming models to reduce inference latency.
AutoBaxBuilder automates creation of security benchmarks for LLM-generated code using bootstrapping techniques.
FusionRoute enables token-level collaboration between specialized LLMs to achieve broad domain performance without expensive scaling.
Training method for efficient LLM distillation addressing performance degradation in student models with strong reasoning ability.
Proposes LEMUR, learned multi-vector retrieval system improving on ColBERT's late interaction model for information retrieval efficiency.
Explores using large language models to make progress on the Gilbert-Pollak Steiner ratio conjecture in computational geometry.
Proposes end-to-end semantic ID generation approach for generative advertisement recommendation systems using alternative to residual quantization.
Proposes SWE-MiniSandbox, container-free reinforcement learning method for scalable training of software engineering agents without storage overhead.
Introduces MoralityGym benchmark with 98 ethical-dilemma problems to evaluate moral alignment in sequential decision-making agents using deontic constraints.
Proposes Volterra signature as explicit feature representation for history-dependent systems, alternative to implicit memory mechanisms in RNNs and transformers.
OmniBehavior benchmark evaluates LLMs as user simulators on real-world long-horizon heterogeneous behavior traces.
Token pruning optimization for DeepSeek-OCR visual-language model reducing inference cost while preserving text fidelity.
Unified self-distillation framework for LLMs without external teachers, addressing free-form trajectory supervision challenges.
Evaluation methodology for prompt injection defenses in LLM tutors, analyzing security-usability-latency trade-offs.
TextSeal watermarking technique for LLMs based on Gumbel-max sampling with zero inference overhead and speculative decoding support.
Framework for AI agents to synthesize enterprise context from multiple sources beyond simple retrieval-based approaches.
Method for calibrating LLMs with semantic-level rewards to improve uncertainty estimation in high-stakes applications.
Linear attention mechanism for autoregressive video diffusion reducing quadratic complexity via cross-frame memory.
Study of autonomous AI agents in supply chains using Beer Game, shows reasoning models outperform humans with 67% cost reduction.
Rule-based reward model for evaluating text-to-image generation alignment, offering interpretability over preference-trained models.
Proposes ECUAS_n metrics for evaluating uncertainty-augmented ML systems with principled evaluation framework.
Framework for systematic corpus-level diagnostics of LLM agent execution traces and failure patterns.
Materialize uses LLM-based coding agents with Claude to find bugs in code and pull requests, sharing implementation considerations and lessons learned.
Self-hosted OpenAI-compatible proxy aggregating free-tier LLM providers into single API with automatic budget routing.
SF Swift talk on production AI coding practices including guardrails, scripts, specialized agents for PRs, and agent reference documentation.
NVSentinel GPU fault detection tool with LLM-powered runbook conversion support via NVSX control plane.
Coding agents like Codex and Claude Code are transforming software engineering. Engineers must shift focus from writing precise code to designing agents that produce correct outcomes.
Discussion: LLMs consistently overestimate implementation time, possibly because training data reflects human timelines rather than AI agent capabilities with Claude/GPT.
OpenAI's CFO discusses financing chip and data center commitments through banks, private equity, and federal backing at WSJ Tech Live event.
Vet tool uses LLM to verify git diffs match intended code changes, with walkthrough of debugging p5.js sketches.
Curated directory of AI teammate products including Perplexity Computer, a multi-model AI agent workspace for research, coding, data analysis, and workflow automation.
Gartner names OpenAI a leader in Enterprise AI Coding Agents quadrant, recognizing Codex for innovation and deployment. Industry recognition announcement.
Talk on fullstack agents and generative UI using AG UI framework.