An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
Reference-free evaluation framework for flowchart image-to-code generation using vision-language models without ground truth.
Reference-free evaluation framework for flowchart image-to-code generation using vision-language models without ground truth.
Curriculum learning framework for distilling chain-of-thought reasoning from large to compact student models via structure-aware masking.
MDL-based framework for layer-wise capacity allocation and pruning in large language models using curvature-aware scoring.
Safe Flow Q-Learning extends offline RL with Hamilton-Jacobi reachability for safety-critical control under constraints.
Phasor Transformer using phase-native unit circle representation to reduce quadratic attention bottleneck in long-context sequence learning.
StateLinFormer: linear-attention transformer with stateful training for long-term memory in navigation tasks.
Study of activation-based persona steering on LLM short-answer generation and automated scoring across architectures.
SVD-based vision token pruning method for efficient Vision-Language Models addressing computational bottlenecks in long sequences.
Peer-Predictive Self-Training (PST) framework enables collaborative self-improvement of language models without external supervision using cross-model aggregation.
Fine-tuned LLM with positive-unlabeled learning for predicting metal-organic framework synthesis scale-up from literature, achieving 93.5% accuracy.
Open-source agentic tutoring framework combining LLMs with personalized feedback, difficulty calibration, and citation-grounded problem solving.
Low-precision FP8 arithmetic optimization techniques for large recommendation models, addressing numerical sensitivity and training efficiency.
Continual learning method for multi-agent LLM systems to maintain and adapt inter-agent communication topologies across evolving task streams.
Criterion-centric approach for pairwise preference prediction in code generation using LLM judges with explicit rubric alignment.
Multi-stage agentic pipeline for generating structured video annotations and Chain-of-Thought reasoning traces for training vision-language models.
Analysis of hallucinations in multimodal LLMs as attention distraction phenomenon, with proposed correction mechanism.
Mechanistic study of RLHF failures including reward hacking and policy collapse, with diagnostic tools and multiple evaluation methods.
Multi-agent digital twin system for autonomous heterogeneous catalyst discovery using machine learning and gas-solid/liquid-solid modeling.
arXiv paper identifying and analyzing temporal preference subgraphs in LLMs through causal localization, showing how models represent temporal tradeoffs internally.
arXiv paper on EasyLens, a training-free method to amplify subtle-lesion representations in medical vision-language models for improved clinical detection.
arXiv paper proposing MeCo, a one-step generative corrector using MeanFlow for multi-channel speech separation with improved perceptual quality.
Addresses modal isolation in multimodal interleaved thinking models where text and vision diverge, proposing stepwise reinforcement for modality supervision.
Benchmark of diffusion policies for robotic manipulation with incrementally increasing context lengths to enable memory and long-horizon task performance.
Hierarchical multi-agent architecture combining LLM-based strategic planning with specialized RL skill policies for complex coordinated decision-making.
Empirical study of Mixture-of-Experts language models on consumer/edge hardware, questioning whether FLOP advantages translate to actual speedups and cost reductions.
Video-SALMONN-R³ implements two-stage video understanding with LLMs: coarse localization then high-fidelity re-watching for efficient question answering.
Play2Perfect investigates pretraining strategies for dexterous multi-finger robotic assembly tasks combining imitation learning and reinforcement learning.
Study on leveraging synthetic TTS speech for training ASR systems in privacy-constrained domains like banking and healthcare, addressing synthetic-real data gaps.
Theoretical characterization of grokking phenomenon using stochastic-geometric analysis of solution space topology and delayed generalization in neural networks.
Diffusion-GR2 uses block-diffusion language models for faster generative reasoning re-ranking with parallel decoding instead of sequential autoregressive inference.
TAG framework for reliable artifact generation with LLMs using test-driven validation, shifting focus from generation quality to validation rigor.
Study analyzing LLM failures in applying Cognitive Behavioral Therapy frameworks despite high theoretical knowledge, highlighting reasoning gaps in practical application.
Approach using local LLMs to extract quantitative data from text for developing fuzzy cognitive maps.
Framework for agentic visual generation that extends knowledge boundaries by searching for out-of-distribution content beyond training data.
Alternative to transformer architecture using causal resonant field mixing for efficient long-context language modeling.
Parameter-efficient continual fine-tuning framework for LLMs using spectrum-aware recursive consolidation of LoRA adapters.
Method for extending LLM context length beyond pretraining windows using dynamic bifocal RoPE for long-context applications and agentic workflows.
Method for improving code generation in low-resource programming languages using test-time compute and difficulty-based data curation for fine-tuning.
Foundation models for earth observation using satellite data with multimodal pretraining for remote sensing downstream tasks.
Survey of combinatorial optimization approaches for trustworthy ML covering transparency, robustness, fairness, and certifiability.
Hamiltonian video dynamics models enabling variable temporal resolution predictions for hierarchical planning and sim-to-real transfer.
Survey of design paradigms and evaluation practices in deep reinforcement learning research spanning algorithms from DQN to model-free methods.
Theoretical proof of robustness law for two-layer neural networks without weight restrictions, extending Bubeck-Sellke conjecture.
Framework for continual learning in LLMs distinguishing between domain adaptation and competence improvement as world conditions change.
NFTR method for offline goal-conditioned reinforcement learning using normalizing flows to address optimistic bias and mode collapse in subgoal selection.
Physics-informed ML methodology for small-dataset manufacturing using abrasive waterjet milling, addressing data cleaning and curation with physics integration.
Theoretical analysis of gradient descent dynamics in deep scalar linear networks showing optimal learning rate scaling depends on data properties.
Survey of multimodal unlearning methods for VLMs, DMs, LLMs across vision, language, video and audio with datasets and benchmarks.
Latent Personality Alignment method for efficient LLM safety training using 66 harm-agnostic statements instead of large adversarial datasets.
Python package implementing PathBoost gradient boosting for interpretable graph-level predictions via discovered path-based features.