Autonomous Algorithm Discovery for Ptychography via Evolutionary LLM Reasoning
LLM-driven framework autonomously discovers and evolves regularization algorithms for ptychography imaging using evolutionary mechanisms and code generation.
LLM-driven framework autonomously discovers and evolves regularization algorithms for ptychography imaging using evolutionary mechanisms and code generation.
LTLGuard uses compact LLMs with symbolic reasoning to translate natural language requirements into formal LTL specifications.
Theoretical analysis of Best-of-N sampling for inference-time LLM alignment, examining optimality and reward hacking vulnerabilities.
TML-Bench evaluates 10 open-source LLMs as autonomous data science agents on Kaggle tasks with time constraints and repeated runs.
Subspace-aware model merging technique for domain generalization across multiple task-specific fine-tuned models.
Depth Charge jailbreak attack targets deep safety attention heads in open-source LLMs through internal model components.
Disentangled Safety Hypothesis reveals LLMs separate harmfulness recognition from refusal behavior, explaining jailbreak vulnerability.
PVminerLLM extracts structured patient voice data from patient-generated text using LLMs for healthcare outcomes research.
Study of LLM-generated nudges and dual-calibration for diverse news recommendation in real-user testing with 120 readers.
BM25-V applies BM25 scoring to sparse visual-word activations from SAE on Vision Transformer features for interpretable image retrieval.
Proof-of-guardrail cryptographic system enabling AI agent developers to prove safety measures are enforced using open-source guardrails.
StreamWise system for real-time serving of multi-modal generative models at scale with efficient model coordination.
Taxonomy of epistemic risks from LLMs collapsing ambiguous concepts in content moderation, hiring, and constitutional AI alignment.
MaCS regularization framework for improving calibration and robustness of deep vision classifiers under distribution shift.
Lexara toolkit for developer evaluation of LLMs in conversational visual analytics, designed with input from 22 CVA developers and 16 end-users.
White-box analysis of how LLMs internally represent and reason about trust using contrastive prompting on gpt-j-6B.
Ensemble learning approach combining CNNs and Vision Transformers for remote sensing image classification.
Survey of foundation models and AI agents in computational pathology, examining clinical integration challenges and real-world adoption barriers.
arXiv paper analyzing consistency bugs in long-form story generation by LLMs, introducing benchmark for narrative consistency.
arXiv paper on using LLM reasoning with reference-guided policy optimization for molecular structure optimization.
arXiv paper on LUMINA using LLM reasoning to guide GPU architecture design space exploration for LLM inference.
arXiv paper on CORE-Seg using reinforcement learning with MLLMs for complex medical lesion segmentation.
arXiv paper on BlackMirror framework for detecting backdoored text-to-image models in black-box settings.
arXiv paper on RAC, a Rectified Flow Auto Coder replacing VAE with multi-step decoding and bidirectional inference.
arXiv paper addressing ecological fallacy in larger LMs by modeling author language context to improve performance.
arXiv paper on XAI approach transforming execution traces from LLM-based coding agents into actionable debugging insights.
arXiv paper on E-AdaPrune, an energy-driven adaptive token pruning framework for efficient Vision-Language Models.
arXiv paper using LLMs to predict well-being and identify self-states in longitudinal social media data based on psychological theory.
arXiv paper on DMM framework for data-free model merging across domains while preserving privacy and avoiding retraining.
arXiv paper on enabling skeleton representation learning by encoding 3D skeleton data for use with vision-pretrained models.
arXiv paper on ProCap framework for change captioning using dynamic procedure modeling instead of static image comparison.
Reinforcement learning approach for off-road autonomous driving addressing long-horizon planning and adaptable control in unmapped, variable terrain.
Multimodal framework leveraging vision-text LLMs for irregularly sampled time series forecasting with contextual semantics and temporal patterns.
Train-free attention recalibration method improving Vision-Language-Action model reliability on out-of-distribution language instructions for robotic manipulation.
RepKAN architecture combining CNNs with Kolmogorov-Arnold Networks for interpretable remote sensing image classification via spectral fingerprinting.
Framework for orchestrating LLM-based multi-agent systems using graph-centric workflow modeling and vibe graphing to reduce implementation complexity.
Retrieval-augmented approach for intent clarification in conversational search systems informed by sensitivity analysis in exploratory search.
Examines intermediate activations in lightweight vision-language models to understand failure modes on visual reasoning tasks relevant to autonomous driving.
Investigates open-weight LLMs for automated essay scoring on Austrian A-Level German essays to reduce grading workload and subjective bias.
Formalizes lifelong embodied navigation learning where LLM-powered agents continually acquire navigation skills across multiple scenes while avoiding catastrophic forgetting.
Studies how domain-specific continued pre-training shapes LLM personality traits and influences problem-solving performance through diverse behavioral tendencies.
Proposes partial policy gradients approach for reinforcement learning in LLMs by optimizing subsets of future rewards for more reliable policy learning.
MLLM framework for physically consistent video object insertion using environment-aware reasoning from multimodal large language models.
Proves predictive coding graphs are a mathematical superset of feedforward neural networks, bridging neuroscience-inspired models with contemporary ML.
VLM-RobustBench: comprehensive benchmark evaluating robustness of vision-language models across 133 corrupted settings and perturbation types.
RAPTOR: controlled study of compact SSL backbones for audio deepfake detection using pairwise-gated transformer fusion.
Inference-time enhancement strategies for flow-matching text-to-image diffusion models to improve generation quality.
CRIMSON: clinically-grounded LLM-based evaluation metric for chest X-ray report generation assessing diagnostic correctness and safety.
Whisper-CD: training-free contrastive decoding framework reducing hallucinations in long-form speech recognition.
MAPO: reinforcement learning approach for optimizing long-horizon multi-turn dialogue policies with outcome-only supervision.