DTBench: A Synthetic Benchmark for Document-to-Table Extraction
Benchmark dataset and evaluation of LLM performance on document-to-table extraction with complex reasoning requirements.
Benchmark dataset and evaluation of LLM performance on document-to-table extraction with complex reasoning requirements.
Empirical study of LLM capability to generate and verify formal ACSL specifications for C programs automatically.
Generative Speech Reward Model for evaluating and improving naturalness in speech language model outputs via RLHF.
Empirical comparison of AI agent-driven social network topology (Moltbook) versus human-driven (Reddit).
GREPO benchmark for GNNs on repository-level bug localization, addressing context limitations of standard LLMs.
End-to-end learned tokenization for LLMs using RL to optimize token boundaries instead of fixed hardcoded compression.
Experiential RL training paradigm embedding explanatory chains for LMs to learn from sparse delayed environmental feedback.
Eureka-Audio: 1.7B parameter compact audio language model matching performance of 7B-30B models on ASR and audio understanding.
NP-specific chemical language models using Mamba state-space models for natural product property prediction and generation.
WoVR: framework using learned world models as simulators for RL-based post-training of Vision-Language-Action policies.
Information-theoretic analysis of trade-off between explanation sufficiency and conciseness in LLM chain-of-thought reasoning.
Comprehensive study on LLM post-training pipeline (SFT, preference optimization) applied to vulnerability detection in code.
EIDOS foundation model family for time series using latent-space predictive learning instead of direct future prediction.
LLM-based social simulation framework modeling dynamic group-level value evolution over time rather than static snapshots.
Adapts LLaVA vision-language model framework to Polish language using automated pipeline for annotation-efficient multilingual VLM development.
Reinforcement learning approach using policy gradients with entropy annealing for parameter-efficient fine-tuning of vision models preventing catastrophic forgetting.
Framework analyzing LLM factuality by distinguishing knowledge encoding failures from retrieval failures at the fact level.
TabTracer applies Monte Carlo Tree Search with LLMs for complex table reasoning, enabling step-level verification and backtracking to reduce errors.
Uses LLMs to predict adversary behavior and anticipate cyberattacks in DevSecOps cloud environments.
Multi-scale agentic AI framework for autonomous O-RAN network control and management, handling complex multi-layer control loops in 6G systems.
DenseMLLM enables multimodal LLMs to perform dense prediction tasks like semantic segmentation and depth estimation without task-specific decoders.
Multi-agent framework combining fine-tuned GPT, LLaMA, and DeepSeek R1 with evidence retrieval and bias detection for medical question answering.
Introduces Pivot-Driven Resampling for exploration in LLM reinforcement learning, improving trajectory discovery within limited sampling budgets.
Proposes UniWeTok, a unified discrete tokenizer with 2^128 codebook for multimodal LLMs enabling high-fidelity reconstruction, semantic extraction, and generation.
Evaluates context window utilization across LLMs including GPT-5, comparing theoretical capacity vs practical performance on long-context tasks requiring detailed understanding.
Framework for LLMs to abstain from answering uncertain scientific claims, decomposing claims into verifiable conditions.
Multimodal framework enabling adaptive visual focusing in vision-language models for ultra-high-resolution remote sensing analysis.
Attack methodology demonstrating skill-based prompt injection vulnerability in coding agents via trace-driven refinement.
Multi-stage workflow using reasoning language models to assess parental cooperation in child protection case reports.
Analysis identifying five recurring biases in financial LLM evaluations: look-ahead, survivorship, narrative, objective, and cost.
Optimization technique for reducing KV-cache memory overhead in vision-language models processing long-form video.
Hybrid GNN model combining temporal graph networks with graph structure learning for dynamic link prediction.
Multi-agent debate framework treating model disagreements as signals for improved tabular anomaly detection.
Benchmark for evaluating LLM agents on real-world advertising analytics tasks with multi-round trajectory-aware workflows.
Epistemological analysis of sycophancy in LLMs and its effects on user beliefs and worldview formation.
Framework using transformer language models to perform causal inference from unstructured text data.
Testing methodology framework addressing challenges in validating AI/ML and quantum computing systems.
Adaptive multi-turn LLM-based framework for optimizing respondent selection in information elicitation surveys.
Multimodal dataset for automated scholarly peer review research based on F1000Research platform data.
Agentic RL framework using memory augmentation for continual CUDA code optimization across GPU architectures.
Design pattern for integrating pre-trained ML models as callable tools within LLM agent workflows for hybrid reasoning.
Large-scale empirical study examining whether AI agent societies develop social convergence dynamics similar to human systems.
Federated learning approach for training Mixture-of-Experts LLMs using heterogeneous edge devices via knowledge distillation.
Adaptive efficient rollout optimization for GRPO-based LLM post-training reducing computational waste when rollout groups share outcomes.
Approach for zero-shot instruction following in multi-task RL using linear temporal logic representations for temporally extended tasks.
AXE agentic system automatically exploits and validates zero-day vulnerability reports using vulnerability metadata and source code.
Ethnographic study examining challenges in designing and evaluating LLMs with domain experts through 12-week pedagogical chatbot development.
Safety audit of Clawdbot, a self-hosted tool-using AI agent, evaluating risks across six dimensions using trajectory-based evaluation.
InnoEval framework for evaluating research ideas using LLMs with knowledge grounding and multi-perspective reasoning across evaluation dimensions.
Methods for differentially private retrieval-augmented generation protecting sensitive data while reducing hallucinations in LLM outputs.