Visual Prompt Discovery via Semantic Exploration
Method for generating visual prompts via semantic exploration to improve LLM perception and visual reasoning capabilities.
Method for generating visual prompts via semantic exploration to improve LLM perception and visual reasoning capabilities.
SlideFormer system for efficient LLM fine-tuning on single GPUs using asynchronous computation overlapping and CPU-GPU co-design.
Study of vision-language model robustness under distribution shifts in visual deductive reasoning tasks.
Research on dexterous robotic manipulation using sparse taxonomy guidance for grasp planning with reinforcement learning and multi-finger control.
Optimization method for tree-based speculative decoding that balances computational overhead with token acceptance in LLM generation, improving inference efficiency.
Evaluation framework for retrieval-augmented generation systems on high-redundancy corpora like financial reports and legal documents, addressing real-world RAG limitations.
arXiv paper on TaNOS framework improving numerical reasoning over tables via operation sketches and self-supervised learning.
arXiv paper analyzing LLM prompt sensitivity, showing shared lexical task representations explain behavioral variability across prompting styles.
arXiv paper using LLMs to guide neural architecture search by translating architectural knowledge into executable code edits.
arXiv paper on KV-cache optimization for static-graph LLM serving. Infrastructure-focused, limited LLM application relevance.
arXiv paper on executable benchmarking suite standardizing evaluation of tool-using agents across web, code, and task environments.
arXiv paper on safety mechanisms for tool-composing agents via monotonic capability attenuation to prevent unsafe end-to-end effects.
arXiv paper showing knowledge graphs improve LLM agent accuracy in industrial asset operations vs flat document stores.
arXiv benchmark dataset (BenGER) for evaluating LLM reasoning on German legal subsumption tasks.
arXiv paper presenting INFUSER, iterative co-training framework for self-improving LLM reasoning without heavy supervision.
arXiv paper on credit distillation method for training long-horizon tool-use agents with reinforcement learning.
arXiv paper proposing COM-as-Action paradigm for AI agents manipulating professional software via component objects.
arXiv paper studying security of agentic browsers, examining same-origin policy effectiveness with AI agents.
arXiv paper proposing inference-time scaling method for LLM reasoning via deterministic layer recursion.
Shows protein contact prediction signals are concentrated in attention heads of language models, enabling efficient single-pass prediction.
RigorBench evaluates autonomous coding agents beyond correctness, measuring engineering discipline and process quality in problem-solving approaches.
Uses contrastive logit steering to identify linear refusal mechanisms in safety-aligned LLMs, showing safety compliance as manipulable linear feature.
Studies over-alignment issues in multilingual LLMs for criminal law translation and summarization in Swiss courts.
Evaluates safety guardrails of eight LLMs across 16 psychiatric conditions using adversarial attacks and introduces harm taxonomy framework.
Improves language model training when using iterate-averaged models by reformulating optimizer design as optimal control problem.
Reduces VLM inference costs by curating pretraining data to induce concise outputs rather than using traditional model compression techniques.
Reflect-R1 improves long video understanding through evidence-driven self-correction using external verification rather than internal reflection.
SHARD protects dense vector stores from inversion attacks by using cell-keyed residual splitting to resist alignment-based attacks on RAG systems.
Studies how inference speed optimizations in embodied AI models affect action quality differently than static ML tasks through empirical analysis.
Generates diverse, realistic bugs in code using LLM-based diffs to create larger-scale, more representative bugfix benchmarks for evaluating LLM capabilities.
TraceLab provides a public dataset and analysis of real coding agent workloads across multiple LLM-based agents and models to improve serving efficiency.
Training LLMs to generate executable workflows as zero-shot structured solutions encoding task-level algorithmic patterns for reproducible deployment.
Analysis of why deterministic few-step generation fails for text latents due to geometric inability to resolve discrete choices before sharp readouts.
Hierarchical Global Attention drop-in replacement for dense causal attention enabling 64K-token context without retraining or parameter changes.
Process sidecar method for revoking learned memory in adapted language models through two-coefficient edits after safety fine-tuning.
First-principles reduced-order model of GRPO training dynamics for LLM reasoning with closed-form analysis replacing empirical hyperparameter tuning.
Depth-wise gradient augmentation optimization paradigm for transformers leveraging structured layer relationships to improve training dynamics.
Imitation learning theory explaining why on-policy distillation outperforms offline supervised fine-tuning under noisy expert feedback using theoretical model.
Quality-aware modulation for diffusion transformers improving image generation by incorporating quality information into denoising process modulation.
Training-free diffusion model acceleration using optimal transport for geometry-aware caching schedule prediction across inference budgets.
LLM-based clinical decision support for epilepsy medication prediction in resource-constrained settings, adapted to local practice with deferral capability.
Knowledge distillation from reasoning model to compact student via dual-agent Chain-of-Thought generation, fine-tuned with LoRA on edge hardware.
Fine-tuning method protecting LLM capabilities by optimizing activation subspaces rather than parameter distances during domain adaptation.
Statistical mechanics framework explaining neural network training dynamics, memorization, and low-dimensional structure of learned weights.
Study of tabular foundation models' generalization to biomolecular property prediction from limited labeled data in few-shot regime.
Benchmark with 120 PowerPoint tasks evaluating computer-use AI agents on multimodal content creation and presentation editing scenarios.
LLM routing system with compliance gates and tiered model selection for regulated industries, optimizing cost and geographic data sovereignty.
Unified offline reinforcement learning framework for optimizing nonlinear multi-objective preferences balancing risk and fairness.
Framework probing parametric memorization in tabular foundation models using in-context learning, separating context-based from memorized predictions.
Web-based toolkit for synthetic tabular data generation using Bayesian mixtures, diffusion models, and latent-space methods with privacy preservation.