Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
AwaRes framework for efficient vision-language models using spatial-on-demand high-resolution crop retrieval to balance accuracy and computational cost.
AwaRes framework for efficient vision-language models using spatial-on-demand high-resolution crop retrieval to balance accuracy and computational cost.
Multimodal LLM for agriculture with Vision-to-Verified-Knowledge pipeline addressing lack of domain-specific agricultural datasets and expertise.
Aegis architecture for cryptographic runtime governance of autonomous AI systems, enforcing policy constraints as verifiable execution conditions.
GridReg framework investigating optimal sparse control-point parameterization for learning-based medical image registration.
Omni IIE Bench, a high-quality benchmark for evaluating instruction-based image editing models across tasks of varying semantic scales.
Optimization techniques for storage and loading of 3D point cloud data in computer vision and deep learning applications.
EmergeNav framework for zero-shot vision-and-language navigation using structured embodied inference to improve VLM execution in continuous environments.
Survey of deployment constraints and mitigation strategies for foundation models in edge embodied systems with real-time control requirements.
PhysQuantAgent inference pipeline for mass estimation in vision-language models to improve robotic manipulation by enabling physical property reasoning.
Machine learning framework for optimizing 2D dendrite synthesis via active learning, synthesis automation, and mechanism discovery.
Adversarial robustness evaluation of open-source VLMs (LLaVA, Qwen2.5-VL) against gradient-based attacks in e-commerce environments.
CineSRD system for speaker diarization in films and TV series using visual, acoustic, and linguistic cues in open-world scenarios.
MSRAMIE training-free agent framework for multi-instruction image editing using multimodal foundation models without additional dataset annotation.
DeepStage framework using deep reinforcement learning and graph neural networks for autonomous defense against multi-stage APT campaigns.
Method for continuous learning in multimodal egocentric activity recognition using modality-aware novelty detection in non-stationary streams.
Beyond Images pipeline for automatic enrichment of multi-modal knowledge graphs by retrieving and converting images to text descriptions.
Survey of 65 developers examining generative AI impact on software development lifecycle tasks, showing 70% reduction in time for design, implementation, testing, documentation.
TorchNWP compiler library tool for coupling AI components with numerical models, improving cross-language compatibility and data transfer efficiency.
Empirical analysis and optimization recipes for efficient deployment of vision-language models in resource-constrained environments, addressing inference bottlenecks.
Robustness evaluation benchmark for NL2SQL systems under perturbations, comparing traditional and agentic LLM settings in dynamic database environments.
Data synthesis approach for improving vision-language model reasoning through multi-hop chains, addressing perception and hallucination errors.
Method for targeted sound detection using shared representation learning and conditional embedding vectors for audio source separation.
Proposes dependence fidelity as evaluation criterion for generative models to assess preservation of multivariate dependence structures in synthetic data.
Systematic study of DPO alignment applied to unified multimodal models, finding generation quality resists alignment across tested conditions.
Study of vector quantization behavior in tokenization for LLMs and generative models, investigating codebook diversity issues and proposing fixes.
Research on evaluating LLMs on ill-defined tasks with ambiguous criteria, analyzing why existing benchmarks fail for complex instruction following.
Analysis of large reasoning LLMs showing cross-lingual knowledge transfer gaps are primarily script barriers rather than language or family effects.
RL agent trained with AlphaZero-style loop to discover efficient arithmetic circuits for computing polynomials using addition and multiplication gates.
Framework for localizing parametric knowledge in mixture-of-experts LLMs using cross-lingual inconsistency as interpretability signal.
SLUMP benchmark measures faithfulness loss when coding agent specifications emerge through interaction, tracking design commitments across long coding sessions.
Study of family bias in Vision-Language Model ensembles, showing models from same family share correlated errors that reduce effective ensemble diversity.
Comprehensive security assessment comparing vulnerability of major LLM architectures to adversarial attacks, providing defensive framework for secure deployment.
REAL method applies regression-aware RL to train LLMs-as-judges that assign numeric scores, preserving ordinal structure in evaluation tasks.
Position paper arguing intent formalization—translating informal natural language requirements to precise specifications—is critical for reliable AI-generated code.
PAuth authorization framework enables precise task-scoped permissions for AI agents interacting with web services, replacing operator-scoped OAuth models.
Study showing generalist multimodal LLMs can gain biometric expertise for iris presentation attack detection via human salience training.
Black-box scanning approach detects data poisoning attacks in code generation LLMs by checking for vulnerability-inducing code generation patterns.
Unsupervised detection of adversarial documents in retrieval augmented generation systems to defend against context manipulation attacks.
Activation probing technique detects motivated reasoning in LLMs where chain-of-thought rationalizes answers influenced by injected hints without acknowledging them.
OPERA data pruning framework improves efficiency and effectiveness of dense retriever finetuning by exploiting heterogeneous training pair quality.
Framework for adaptive contracts in AI delegation that selectively performs detailed evaluation to balance noise reduction against evaluation costs.
LLM-driven on-premise pipeline for anonymizing text by replacing PII with realistic type-consistent surrogates while preserving data utility.
Empirical study comparing 120 base-aligned LLM pairs on 10K human decisions, showing aligned models underperform at predicting actual human behavior in strategic games.
TharuChat applies synthetic data generation and human validation to bootstrap LLMs for Tharu, a low-resource Indo-Aryan language spoken by 1.7M people.
Analysis of spatial understanding capabilities in multimodal LLMs for segmentation tasks via layerwise probing and attention mechanisms.
Research on low-bit quantization techniques for Kolmogorov-Arnold Networks to enable efficient inference.
Study deploying LLM-based tool integrated with EHR system to automate surgical patient triage at Stanford Health Care.
DANCE dynamically prunes 3D CNNs at frame, channel, and feature levels to maximize energy efficiency for edge video processing.
GUIDE is open courseware with runnable Colab labs teaching generative AI with standardized units of slides, videos, labs, and papers.
Multi-agent MLLM system for long-form video understanding using cognitive inspiration and task decomposition with improved context management.