ReLMXEL: Adaptive RL-Based Memory Controller with Explainable Energy and Latency Optimization
Multi-agent reinforcement learning framework for dynamically optimizing memory controller parameters with explainable energy and latency tradeoffs.
Multi-agent reinforcement learning framework for dynamically optimizing memory controller parameters with explainable energy and latency tradeoffs.
Vision-language model method for embodied agents to estimate long-horizon task progress using recurrent reasoning with efficient video processing.
Diffusion-based approach for learning probability distributions over permutations using reflected diffusion on the symmetric group.
WebPII benchmark with 44,865 annotated e-commerce images for detecting personally identifiable information in web screenshots for privacy-preserving computer-use agents.
Analyzes internal representation shifts during VLM jailbreaks and proposes defense mechanisms based on distinguishing benign from harmful inputs in representation space.
Online learning algorithm improving data efficiency of RLHF by incrementally updating reward and language models during preference learning.
SCALE predicts cellular responses to perturbations using scalable transport models for virtual cell experimentation from single-cell measurements.
CRE-T1 goes beyond contrastive learning for reasoning-intensive retrieval by dynamically identifying implicit reasoning relationships between queries and documents.
Security framework for autonomous LLM agents addressing vulnerabilities including unauthorized instruction compliance, information disclosure, and identity spoofing in healthcare deployment.
Phasor Transformer reduces attention bottlenecks in transformer models for long-context sequences using phase-native representations on the unit circle manifold.
Baguan-TS combines sequence representation learning with in-context learning for time series forecasting using 3D Transformers, enabling gradient-free adaptation.
AdaZoom-GUI improves vision-language models for GUI grounding by using adaptive zoom to handle high-resolution screenshots and ambiguous instructions for UI automation.
VLM2Rec investigates vision-language models as multimodal encoders for sequential recommendation systems.
Cross-attention mechanism using beneficial noise for unsupervised domain adaptation in transformers.
UniSAFE benchmark for system-level safety evaluation of unified multimodal models across 7 modalities and tasks.
AutoML approach combining deep unfolding with automated parameter learning for wireless beamforming optimization.
Sustainable federated learning framework using pre-trained model quantization to reduce energy overhead on IoT devices.
Comprehensive benchmark of AI-generated text detectors across multiple architectures, domains, and adversarial conditions.
Vision-language-action model with kinematic attributes for fine-grained robot manipulation from dense language commands.
Teacher-student framework for multi-class and continual visual anomaly detection in industrial inspection.
SE(3) equivariant point cloud analysis using coordinate-based convolutional kernels with group convolution.
Continual learning technique using adaptive normalization to handle data distribution shifts in dynamic environments.
Explainability framework for tabular foundation models in outlier detection with modular signal generation.
Learning latent actions and environment dynamics from offline trajectories without observed actions using demonstrator diversity.
Goal-regressive planning approach for 3D indoor scene editing from natural language instructions.
Browser extension combining curated dictionary with OpenAI LLM to provide contextual help for technical terms via tooltips.
Neural network training robust to both label noise and adversarial attacks via unified loss function approach.
Benchmarking framework for reinforcement learning algorithms using stochastic converse optimality with known optimal policies.
Evolutionary algorithms to automatically design efficient multigrid cycles for solving partial differential equations.
Cross-domain few-shot learning with vision-language models like CLIP for fine-grained visual recognition tasks such as medical diagnosis.
Benchmark and analysis of hallucinations in multimodal LLMs with fine-grained negative queries covering multi-object, multi-attribute, and multi-relation scenarios.
Post-training framework for small local LLM agents performing Linux privilege escalation tasks with verifiable rewards and resource constraints.
Adaptive guidance method for retrieval-augmented masked diffusion models that resolves conflicts between retrieved context and parametric knowledge.
Benchmark for evaluating vision-language model reasoning and segmentation performance under adverse weather conditions with degraded visual cues.
Framework for testing LLM trading agents with anonymized market data to validate genuine market understanding versus memorized ticker associations.
Evaluation of Segment Anything Model 3 for eye image segmentation comparing text and visual prompting modes against prior SAM versions.
Survey of machine learning methods for network intrusion detection and adversarial learning techniques for synthetic attack data generation.
Training-free fine-grained visual recognition method using sample-wise adaptive reasoning with large vision-language models for subordinate-level category classification.
Analysis showing attention sinks in transformers induce gradient concentration during backpropagation under causal masking, affecting training dynamics.
LLM training approach using generator-verifier co-evolution to escape consensus trap and improve reasoning without ground-truth labels via reinforcement learning.
Theoretical analysis of sparsity in infinite-width shallow ReLU networks trained with total variation regularization using duality theory.
On-model anomaly detection method that leverages primary model representations to detect distributional shifts without separate AD models.
Video world model framework with inverse dynamics rewards that ensures generated robot action sequences satisfy rigid-body and kinematic constraints.
Post-training quantization method for large vision-language models using quantization-aware integrated gradients to reduce memory and computational overhead.
Study of dropout-induced variability and uncertainty in transformer models across 19 architectures using Monte Carlo sampling for inference-time evaluation.
Multimodal framework for automated program repair using LLMs that jointly reasons over code, issue descriptions, and GUI screenshots to improve debugging workflows.
Reinforcement learning method for training code search agents to localize relevant files, classes, and functions in large repositories as prerequisite for coding tasks.
arXiv paper on inferring stage-play spatial layouts from narrative text, demonstrating language model spatial reasoning for automating dramaturgy.
arXiv paper proposing GeCO, time-unconditional flow matching framework for adaptive robotic control using diffusion models.
arXiv paper investigating how LLMs compute verbal confidence scores and whether they're generated just-in-time or cached during inference.