Efficient Autoregressive Inference for Transformer Probabilistic Models
Paper on efficient autoregressive inference for transformer probabilistic models balancing set-conditioning with joint distributions.
Paper on efficient autoregressive inference for transformer probabilistic models balancing set-conditioning with joint distributions.
Paper on test-time prior adaptation for simulation-based inference using diffusion models for Bayesian inference.
TROJail uses trajectory-level optimization with process rewards to learn multi-turn jailbreak strategies against LLMs.
Stream.FM applies flow matching for real-time streamable speech restoration with 32ms algorithmic latency.
DRAM framework combines mechanism design and online learning for sequential multi-agent truthful reporting.
Fitted Q-evaluation theory for off-policy reinforcement learning without requiring Bellman completeness using stationary weighting.
QSLM quantization framework with tiered search optimizes spike-driven language models for embedded deployment.
Study investigates whether LLMs encode functional importance of individual reasoning tokens for reasoning chain compression.
Theoretical analysis of local updates in distributed optimization showing acceleration benefits and topology effects in federated settings.
SAGE-32B is a 32B parameter model fine-tuned via iterative distillation for agentic reasoning, task decomposition, and tool usage.
Multi-Focus Attention Instruction probe disentangles recognition vs synthesis failures in multi-hop LLM reasoning.
Study reveals LLMs are highly sensitive to prompt order in multiple-choice QA due to causal attention limitations.
Temp-R1 is an autonomous agent for temporal knowledge graph question answering trained via reverse curriculum reinforcement learning.
AskBench evaluates and improves LLM ability to request clarification on ambiguous prompts using reinforcement learning with rubric guidance.
CLIPoint3D adapts vision-language models like CLIP for 3D point cloud domain adaptation with few-shot unsupervised learning.
ConFu improves speculative decoding for LLM inference acceleration by enhancing draft model quality to propose better candidate tokens for verification.
Self-distillation degrades LLM reasoning by suppressing epistemic verbalization of uncertainty; controlled experiments isolate degradation mechanisms.
PolarQuant post-training quantization for LLMs uses Hadamard rotation and Gaussian weight distribution for near-lossless compression.
Analysis reveals multilingual language models organize internal representations by orthographic script rather than linguistic structure across language families.
Semantic Intent Fragmentation attack exploits LLM orchestration systems where composed subtasks violate security policy despite individual benignness.
Triadic Suffix Tokenization improves LLM numerical reasoning by partitioning digits into three-digit triads with explicit magnitude markers.
Multi-agent study evaluating whether LLMs can cooperate on resource governance through elected leadership and self-governance mechanisms.
Event Tensor abstraction eliminates kernel launch overheads in LLM inference by fusing operators into persistent kernels handling dynamic shapes.
Cross-domain metacognitive benchmark for LLMs with 524 items across six cognitive domains using human psychometric methodology.
BARD framework bridges autoregressive and diffusion vision-language models via progressive block merging and stage-wise distillation for efficient inference.
Q-SINDy integrates quantum feature maps into sparse identification of nonlinear dynamics, addressing coefficient cannibalization failure mode.
Novel framework integrating Uniform Discrete Diffusion Models with Group Relative Policy Optimization for stable RL training on discrete generative models.
Systematic benchmark comparing cloud and open-source LLMs on System Dynamics tasks: causal loop diagram extraction and interactive coaching.
Incomplete article about arXiv Labs framework; content does not match title about LLM benchmarking.
Aide: customizable Android voice assistant supporting Claude, OpenAI, or OpenAI-compatible endpoints with on-device encryption.
Open source local screen memory tool for Claude and coding agents; OCR and summarization via local AI; Mac-only Swift app.
Prismer: infrastructure layer for long-running AI agents with error recovery, persistent memory, and cross-session learning.
Cloudflare's internal AI engineering stack: MCP servers, agent infrastructure, and iMARS tiger team integration for engineering workflow.
Marketing copy for Meticulous automated testing tool for AI-generated code; claims to eliminate debugging.
MemFactory: unified framework for training and inference of memory-augmented LLMs using reinforcement learning for agent memory operations like extraction and retrieval.
MemFactory: unified framework for inference and training of memory-augmented LLMs for long-term AI agents using RL optimization.
Technical article on building search engines for AI agents, handling complex query graphs and agent-generated code constraints.
ASCEND: DevSecOps framework with AI-powered merge conflict resolution integrated into CI/CD pipelines.
Meta installing tracking software (MCI) on employee computers to capture interactions for AI agent training.
Kuri: Zig-based browser automation tool for AI agents with 464KB binary, 3ms cold start, 16% token efficiency improvement.
Benchmark study measuring vulnerability patterns in code generated by AI models under time pressure conditions.
Google WeatherNext 2: state-of-the-art ML models for weather forecasting released to researchers and enterprises.
PayClaw: tool enabling AI agents to manage wallets and execute financial transactions autonomously.
Visualization tool for Group Relative Policy Optimization, an LLM training method.
Voxyflow: personal AI assistant agent that plans, codes, and ships projects. Open source, runs locally. Alpha stage.
Ravix: autonomous AI agent running on Claude Code subscription, auto-manages email inbox. 60-second setup, no additional costs.
AI-powered startup profile submission tool using Claude/ChatGPT agents with MCP integration for form filling and validation.
Grafana Cloud CLI (gcx) enabling AI agents to query production observability without leaving editor.
OpenAI releases open-weight Privacy Filter model for detecting and redacting PII in text. Infrastructure tool for developers building AI applications with privacy protections.
Analysis of GPU cluster costs for AI/ML companies and spending breakdown for foundation models.