GodHands – Deterministic Desktop Automation via MCP
GodHands: MCP server for deterministic desktop automation treating apps as structured APIs instead of vision-based screenshots for high-reliability AI agents.
GodHands: MCP server for deterministic desktop automation treating apps as structured APIs instead of vision-based screenshots for high-reliability AI agents.
Frankenstein-style analysis framework isolating specific visual reasoning improvements from RL versus supervised fine-tuning in vision-language models.
Moltis: Rust-based AI agent with memory, tools, and runtime self-extending skills. Single binary with web UI for local, trustworthy AI assistant.
T3D proposes trajectory self-distillation framework to enable fast parallel token decoding in diffusion LLMs with fewer refinement steps while maintaining generation quality.
DeepGen 1.0 is a lightweight 5B unified model for image generation and editing using Stacked Channel Bridging, achieving competitive performance to larger models with reduced deployment costs.
ExStrucTiny benchmark evaluates VLM performance on schema-variable structured information extraction from diverse enterprise documents with flexible schemas.
Sci-CoE framework enables LLMs to self-improve on scientific reasoning through co-evolution as both solver and verifier with geometric consensus mechanisms.
dVoting fast voting technique for diffusion LLMs enabling parallel test-time scaling for improved reasoning performance.
Theoretical and empirical analysis of on-policy distillation as dense KL-constrained RL, proposing generalized reward extrapolation.
P-GenRM enables personalized LLM alignment through scenario-specific reward models with test-time user-based scaling, addressing generalization to new users with limited feedback.
StateLM foundation model framework giving LLMs agency to manage their own context and memory via database operations.
GigaBrain-0.5M VLA model trained via world model-based reinforcement learning for improved multi-step action prediction.
DeepSight is a unified toolkit for LLM/MLLM safety covering workflow, evaluation, diagnosis, and alignment with integrated explainability and risk scenario grounding capabilities.
LawThinker autonomous legal research agent using Explore-Verify-Memorize strategy with intermediate step verification in dynamic environments.
Composition-RL optimizes RLVR training by composing verifiable prompts to balance hard and easy examples, mitigating ineffective data and enabling better prompt dataset expansion.
Gaia2 benchmark for evaluating LLM agents in realistic, asynchronous, dynamic environments with temporal constraints and collaboration.
TADA: activation steering technique for audio diffusion models to control semantic musical concepts through shared attention layers.
Adaptive framework for intelligent AI agent delegation across decomposed sub-tasks with dynamic adaptation to environmental changes and failure handling.
Region-to-Image Distillation: reduces latency in multimodal LLMs' fine-grained perception by distilling zooming behavior into inference-time efficiency.
Detection method for identifying RLVR training data contamination via structural convergence signatures in reasoning trajectories, addressing benchmark contamination concerns.
GPT-5.3-Codex-Spark real-time coding model with 15x faster generation, 128k context, available in research preview.
Light4D: training-free framework for 4D video relighting under extreme viewpoints using diffusion models with temporal consistency.
MiniCPM-SALA hybrid sparse-linear attention architecture for efficient long-context LLM processing in 9B parameter model.
Code2Worlds: extends coding LLMs to 4D world generation with dynamic physics simulation by addressing multi-scale context and semantic-physical execution gaps.
Framework enabling LLMs to perform in-context exploration through length-incentivized RL, allowing models to generate, verify, and refine multiple reasoning hypotheses within continuous context.
Large-scale study on adapting general-purpose VLMs to e-commerce attribute understanding while preserving generalizability across multi-image noisy product data.
Diffusion language models tailored for CUDA kernel code generation, addressing data scarcity and specialization challenges through parallel token generation approach.
ThinkRouter framework routing reasoning between latent and discrete spaces to improve efficiency based on model confidence dynamics.
ScalSelect training-free data selection method for efficient visual instruction tuning of vision-language models.
ABot-N0 unified Vision-Language-Action foundation model for embodied robot navigation across five core tasks using hierarchical brain-action architecture.
SPES enables memory-efficient decentralized LLM pretraining using mixture-of-experts and distributed GPUs without full model replication on each node.
INTENT framework for budget-constrained LLM agents solving multi-step tasks under monetary constraints via intention-based planning.
MuRGAt benchmark evaluates multimodal LLM attribution and factual grounding across complex reasoning tasks involving multiple modalities and information sources.
MolmoSpaces open ecosystem for large-scale benchmarking of robot navigation and manipulation policies with diverse scenarios.
GameDevBench evaluation framework tests multimodal AI agent capabilities on game development tasks combining code, shaders, sprites, and animations.
ABot-M0: VLA framework with action manifold learning for robotic manipulation, includes data curation pipeline for heterogeneous embodiment data.
MOSS-Audio-Tokenizer proposes end-to-end discrete audio tokenization using homogeneous architectures for improved reconstruction and scaling in audio foundation models.
A2A framework enables connecting AI agents across different frameworks and teams for interoperability.
Research study on attention masking strategies in decoder-only LLMs for user representation learning using contrastive learning on large-scale behavioral data.
Neural Additive Experts framework balancing interpretability and accuracy in generalized additive models through feature interaction gating.
MetaphorStar uses end-to-end visual RL to improve MLLM understanding of metaphorical content in images, enabling multi-hop reasoning and cultural context awareness.
Feature Activation Coverage (FAC): measures post-training data diversity in LLM feature space for more effective downstream task performance.
Contamination-free medical benchmark for evaluating LLMs with automated rubric evaluation, addressing data leakage and temporal misalignment in clinical settings.
Method for reasoning in continuous latent space rather than discrete tokens, addressing feature collapse in latent reasoning paradigms for LLMs.
Theoretical and empirical analysis of safety alignment challenges in self-evolving multi-agent LLM systems; identifies self-evolution trilemma.
χ0 identifies distributional shift across human demonstrations, policy inductive bias, and test-time execution as robustness bottleneck in robotic manipulation, proposing alignment approach.
StealthRL reinforcement learning framework stress-tests AI text detector robustness via adversarial paraphrasing using GRPO and LoRA adapters.
OneVision-Encoder: codec-aligned sparsity principle for multimodal architectures that process sparse discriminative information efficiently.
Curriculum learning approach using code generation for agents to progressively learn in open-ended environments with foundation models.
MemFly: memory optimization framework using information bottleneck principles for LLM agents to balance compression and retrieval precision in long-term memory.