Improving Automatic Summarization of Radiology Reports through Mid-Training of Large Language Models
Proposes mid-training adaptation strategy for LLMs to improve automatic summarization of radiology reports in clinical domain.
Proposes mid-training adaptation strategy for LLMs to improve automatic summarization of radiology reports in clinical domain.
Enhances automated short answer grading using GraphRAG to improve LLM reasoning through structured knowledge representation instead of flat vector retrieval.
Proposes HypeLoRA, a hyper-network framework for parameter-efficient fine-tuning of Transformers, addressing miscalibration in language models via LoRA adaptation.
Evaluates generative AI and LLMs for scoring constructed responses in high-stakes testing, comparing feature-based models to neural approaches.
URAG: benchmark for evaluating uncertainty quantification in retrieval-augmented generation systems across multiple domains.
Analysis of how prompt framing biases LLM decisions in threshold voting tasks with individual-group interest conflicts.
CDEoH: LLM-driven automated algorithm design with category-driven diversity to improve evolutionary stability and convergence.
Stock price prediction approach integrating LLMs with financial news using stock name embeddings for relevance filtering.
Expert prefetching scheme for accelerating MoE model inference by predicting and preloading expert weights in memory-constrained settings.
Review of automatic collaboration analysis methods using task-oriented conversational data for studying human cooperation.
LLM-MRD: distillation approach for multimodal fake news detection combining LLM reasoning with efficient student models.
Self-improvement framework for LLMs using mutual information maximization between user context and responses without additional labeled data.
STEU: parameter-efficient method for unlearning sensitive information from clinical language models while preserving utility.
Study comparing LLM, human, and author assessments of STROBE reporting checklists in observational rheumatology studies.
VERDICT: multi-agent LLM system for interpretable legal judgment prediction with verifiable reasoning across law articles and penalties.
MemReward: graph-based experience memory system for LLM reward prediction with limited labeled data using prior experience.
LeWorldModel: stable joint-embedding predictive architecture for learning world models from raw pixels with minimal loss terms.
Memory-driven role-playing paradigm for LLMs using internal memory stores to maintain consistent personas during long dialogues.
Theoretical framework connecting neural networks with ternary gamma semiring algebra, demonstrating improved compositional generalization.
Low-resource entity matching method using prompt-tuning with attribute guidance, reducing need for large labeled datasets.
Hierarchical proof search framework using LLMs for automated code verification in Lean 4, decomposing complex verification goals into subgoals.
AI-based hardware performance projection system for SoC benchmarking as faster alternative to cycle-accurate simulators.
Framework using LLMs with evolutionary tuning for RTL code optimization prioritizing power reduction while ensuring functional correctness.
Unified framework evaluating 51 post-training alignment algorithms across model scales, revealing scale-dependent performance ranking inversions.
Federated learning approach using diffusion guidance to handle non-IID multimodal data and semantic discrepancies across clients.
Method for compressing embeddings in dense retrieval systems using spectral tempering to balance variance preservation and noise reduction.
Framework for dynamically routing queries to optimal LLMs from large model pools using fine-grained latent task discovery without manual taxonomies.
Study on teaching users privacy protection through integrated tools within conversational agents via in-context experiential learning.
Analysis of capability-alignment paradox where defense training against prompt injection attacks degrades LLM agent performance on autonomous tool use tasks.
Research investigating whether linear probes detect genuine evaluation awareness in LLMs or merely surface-level format sensitivity in benchmarks.
Study on how vocabulary and word order affect learnability in transformer language models across different languages using synthetic variants.
LoFi: fine-grained location-aware representation learning for chest X-ray retrieval and phrase grounding without region-level labels.
TrustFlow: reputation propagation algorithm for multi-agent ecosystems using topic-aware vector reputation instead of scalar scores.
Analysis of multiplicative update fixed-point iteration convergence for nuclear norm optimization in private machine learning.
Framework formalizing security definitions for LLM agents considering contextual action legitimacy and instruction origins.
Adaptive layerwise perturbation method addressing off-policy problems in LLM reinforcement learning training stability.
FedAgain: trust-based federated learning strategy for robust kidney stone identification from endoscopic images across diverse devices.
Gastric-X: multimodal benchmark dataset for training vision-language models on gastric cancer analysis with clinical workflow alignment.
Methods to induce sustained creativity and diversity in LLM outputs for exploratory search tasks requiring iterative discovery.
dinov3.seg implements open-vocabulary semantic segmentation using vision-language models for pixel-level classification of unseen classes.
FDARxBench: benchmark for evaluating LLM document-grounded question-answering on FDA drug labels for regulatory assessment.
Novel kernel framework (UKTL) for learning from higher-order tensor data with efficient similarity measures across tensor modes.
Research on scalar quantization methods for matrix multiplication to minimize mean-squared error, with applications to efficient computation.
PFM-VEPAR uses prompted foundation models for pedestrian attribute recognition combining RGB and event camera data with minimal computational overhead.
Uses graph neural networks to co-design soft robot morphology and control policies, solving the challenge of simultaneous optimization through morphology-aware learning.
Skilled AI Agents addresses hardware-in-the-loop challenges for LLM-based agents in embedded and IoT systems development with timing and peripheral constraints.
Physics-informed neural networks with adaptive clustering predict information cascade popularity using graph convolution and recurrent architectures.
CAF-Score proposes a reference-free metric for evaluating audio captioning by combining CLAP embeddings with large audio-language models.
DeepStock applies policy regularizations from classical inventory theory to deep reinforcement learning for inventory management, improving hyperparameter sensitivity.
MetaCues is an interactive tool using generative AI to promote critical engagement and metacognitive thinking during information seeking tasks.