BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
BaFCo benchmark dataset for Bangla form comprehension using multimodal LLMs, addressing low-resource language document understanding gaps.
BaFCo benchmark dataset for Bangla form comprehension using multimodal LLMs, addressing low-resource language document understanding gaps.
EvalLoop methodology for evaluation-driven iterative improvement of LLM-based business systems, focusing on diagnosis and fixing rather than static model selection.
Empirical taxonomy of code mutations in AI-generated pull requests for performance optimization, analyzing 33k agent PRs.
RPAM: metric for measuring semantic associations and biases in language models with predictive validity for downstream task performance.
Survey study examining human trust in AI-generated legal advice versus human lawyers when advice is legally correct but socially controversial.
SCOReD: chain-of-thought distillation optimization for recommendation systems using student-aware training to avoid verbose reasoning traces.
Survey of execution-layer security research for AI coding agents covering isolation, access control, TOCTOU vulnerabilities, and protocol threats.
Security vulnerability analysis of Model Context Protocol implementations showing tool metadata concealment attacks across AI coding agent servers.
Counterfactual supervision method to train LLMs when to use external search versus parametric knowledge for improved task performance.
Data-dependent evaluation methods for budgeted submodular maximization algorithms in machine learning optimization problems.
Legato 2 pipeline for optical music recognition processing sheet music sequentially to extract symbolic notation and semantic knowledge.
FORGE framework enables robots to generalize tool use across novel objects via keypoint trajectory reasoning for functional transfer learning.
SegAnswer: multimodal LLM method using pixel-level segmentation before answering to improve visual reasoning accuracy in image understanding tasks.
AI-based screening system combining image classification and vessel segmentation for detecting retinopathy of prematurity in infants.
Vulnerability analysis of aligned LLMs comparing same-lineage models to isolate safety behavior from architectural differences in code review terminology.
In-context learning approach for ranking antibody candidates by binding affinity using contextual information from labeled antigen-specific comparisons.
Unsupervised anomaly detection method for identifying information operations users via behavioral and language patterns on social media.
Natural gradient descent variant for differentially private training that accounts for loss curvature to improve optimization efficiency under privacy constraints.
Analytical framework for LLM serving optimization using floor-first residual-driven triage to estimate resource bottlenecks before grid search.
Harrison.Rad 1.5 multimodal foundation model for automated radiology report generation from images, clinical history, and prior studies.
Simple coreset selection method using medoids for efficient few-shot knowledge distillation that surpasses random baseline sample selection.
Selective gradient computation framework reducing training cost by excluding low-loss samples from backward pass with unbiased gradient estimation.
Benchmark and methods for policy-adaptive image safety guardrails that generalize to policy changes without retraining.
Contextual multimodal document retrieval benchmark and methods that preserve textual/visual content while resolving multi-page aggregation queries.
Distributed multi-MCP architecture for vendor-agnostic SDN-based automation and autonomous control of multi-layer IPoDWDM networks with E2E service lifecycle automation.
Cost-efficient influencer matching system using three-stage cascade of small open-weight models instead of frontier LLM prompting for semantic matching on Thai marketing criteria.
Evaluation of LLM-generated metadata for RDF dataset search, comparing rewriting and agentic graph-based generation methods for retrieval effectiveness and faithfulness.
Agentic AI architecture using MCP protocol for autonomous control and lifecycle automation of multi-vendor IPoDWDM networks with closed-loop control validation.
Training-free method to improve grounding confidence in multimodal LLMs by detecting hallucinated spatial/temporal predictions using multi-token localized attention.
PluraMath extends mathematical reasoning evaluation to 99 languages beyond English and Chinese, addressing dataset bias in LLM benchmarks.
Prompt Coach is agentic tutor providing Socratic guidance for developers learning prompt engineering within IDE workflow.
SocaSim is LLM-based multi-agent simulation framework modeling Putnam's Social Capital Theory for studying collective action.
Studies incidental learning loss in AI-assisted software development and proposes design strategies for agents to support developer education.
RoME uses mixture of low-rank experts to achieve robustness against multiple adversarial perturbations with reduced trade-offs.
LLM-guided approach to correct measurement credibility in industrial soft sensing and process inference for robust predictions.
Training-free acceleration method for diffusion and flow matching models via x-prediction without retraining or distillation.
Predicts Java method energy usage incorporating execution time to enable early energy optimization in software development.
Empirical study of fine-tuning and evaluation metrics for neural decompilation of Dart AOT binaries using code-generation models.
Analyzes property-driven synthetic data generation engineering for data-scarce domains like breast cancer treatment planning systems.
Study on LLM agents for deliberative collaboration under partial observability with joint decision-making benchmark and multi-agent coordination.
LongCrafter synthesizes long-context supervised fine-tuning data with evidence graphs to improve LLM understanding across diverse tasks.
X-FEMR proposes token-level explainability approach for Electronic Health Records foundation models using transformers for clinical prediction tasks.
Study investigates reward function design for reinforcement learning to improve LLM-generated BPMN process models beyond supervised fine-tuning limits.
UBEP optimizes communication library for Mixture-of-Experts models on high-bandwidth superpods, addressing serialization and latency bottlenecks.
Spider 2.0-AIFunc benchmark evaluates text-to-SQL models on AI-native SQL workflows with LLM functions for classification, filtering, and search tasks.
UI2App benchmark for evaluating LLMs on visual interaction inference in executable web application generation from screenshots.
Large-scale evaluation of 9 uncertainty estimation methods across 22 languages for LLM multiple-choice question answering.
Framework using LLM code agents for automatic software verification, enabling autonomous proof generation for interactive theorem provers like Coq.
RuBench 1.0 benchmark with 25 repository-level agentic coding tasks in Russian and other languages from live open-source projects.
Experimental design framework for systematically evaluating autonomous model discovery behavior of LLM coding agents with variability quantification.