GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval
Benchmarks GraphRAG with local LLMs on consumer hardware for healthcare EHR retrieval without cloud dependencies.
Benchmarks GraphRAG with local LLMs on consumer hardware for healthcare EHR retrieval without cloud dependencies.
Proves DPO and RLHF equivalence is conditional on implicit assumptions; identifies failure modes and alignment challenges.
DISC uses hypernetworks to generate vision-language manipulation policies that decouple instruction processing from state observation.
PlexRL optimizes cluster-level resource orchestration for reinforcement learning training with verifiable rewards on LLMs.
PlanningBench introduces scalable benchmark generation for evaluating LLM planning capabilities with controllable difficulty levels.
Dynamic Threat Detection Agent uses AI to automatically investigate security incidents and adapt detection logic without human intervention.
RL agent training for fighting games learning variable action durations beyond fixed-frame decisions.
Research on calibration vs decision reliability in machine unlearning for language models.
Analysis of AlltoAll dispatch bottlenecks in mixture-of-experts parallelism across architectures.
Hybrid ML-physics model for forest height estimation from satellite interferometry data.
Concentration bounds for stochastic approximation algorithms under heavy-tailed Markovian noise.
Study comparing persona steering vectors to targeted steering (CAA) for mitigating LLM sycophancy.
Research on conditioning Gaussian processes via diffusion models with closed-form guidance.
Transformer-based mutation operator for Cartesian genetic programming in automated circuit design.
Training multimodal LLMs using paired modalities instead of fully aligned multi-way datasets.
SpectralEarth-FM foundation model integrating hyperspectral imagery with multisensor Earth observation data.
Musical Attention Transformer mechanism to improve music generation quality and reduce repetition.
AIMBio-Mat framework for AI-guided materials discovery and biomedical translation with FAIR principles.
Research on decoupling communication from policy in multi-agent RL under bandwidth constraints.
Byzantine-resilient decentralized federated learning framework (ABC-DFL) for EV battery intelligence.
Research proposes Linear-DPO for aligning diffusion and flow-matching models via preference optimization.
End-to-end autonomous driving using reinforcement learning with cognitive and foresight components to overcome imitation learning limitations.
Automates ICD psychiatric diagnosis classification from Spanish free-text descriptions using NLP and ML techniques on 145,513 clinical records.
Proposes mathematically rigorous, computationally efficient model complexity measure based on gradient similarities across inputs for interpretation and model selection.
RL-based control approach using Y-wise affine neural networks for chemical process management.
Quantum-enhanced reinforcement learning framework for chemical process synthesis with improved scalability.
Federated learning framework for parameter-efficient LoRA fine-tuning of LLMs with heterogeneous clients.
Gradient descent analysis of simplified linear transformer learning in-context regression at large learning rates.
High-fidelity LLM inference simulator supporting disaggregated execution, complex parallelism, and agent workloads.
Prompt optimization method using regularization to prevent distributional overfitting and improve LLM generalization.
Semiparametric debiasing theory for bilevel gradient estimation using efficient influence functions.
Systematic corpus-level trace diagnostics tool for identifying and diagnosing failure patterns in LLM agent execution traces.
Data mixture optimization framework for efficient real-synthetic co-training in autonomous driving end-to-end learning.
Hybrid method combining generative and regression approaches for fast, efficient image restoration via stochastic interpolants.
Theoretical analysis of memorization vs. generalization in diffusion models through independent training on dataset subsets.
Benchmark for standardizing tactile-based reinforcement learning across robotic morphologies with GPU parallelization.
Framework for deciding when to supplement pre-trained simulators with real experiments under budget constraints.
AI-driven platform for publishing and organizing human and AI-generated research, addressing scalability in academic publishing.
Theoretical analysis of GP-UCB optimality for sequential optimization of black-box functions through effective optimism levels.
TRAM enables test-time adaptation of RL agents to new safety constraints by compositing a mixture of pre-trained risk-neutral policies without retraining.
Presents Optimization Hyper-parameter Laws framework for deriving dynamic learning rate schedules and other optimization parameters during LLM training.
Proposes self-improving mechanism for skill-based meta-RL to reduce sensitivity to noisy offline demonstrations in long-horizon environments.
Investigates factors influencing loss-to-loss scaling laws that relate pretraining and downstream task losses for LLM optimization and generalization.
Extends Learning-to-Defer framework to allocate queries to top-k experts instead of single expert, unifying multiple prediction paradigms.
Proposes CT-OT Flow to estimate continuous-time dynamics from discrete temporal snapshots using optimal transport methods.
Introduces GradPower, a lightweight gradient transformation technique using sign-power elementwise operations to accelerate language model pre-training.
Uses heterogeneous prompting techniques to improve LLM performance on time series forecasting tasks compared to traditional deep learning methods.
Proposes Strict Subgoal Execution for hierarchical RL in long-horizon tasks, improving subgoal feasibility and high-level planning reliability.
Presents FAIR-Pruner, a search-free framework for adaptive layer-wise structured pruning of neural networks using within-layer ranking signals.
Addresses spurious correlations in multimodal sentiment analysis using causal attention mechanisms across text, audio, and visual modalities.