ULTRA:Urdu Language Transformer-based Recommendation Architecture
Urdu language transformer-based recommendation system for personalized news retrieval in a low-resource language context.
Urdu language transformer-based recommendation system for personalized news retrieval in a low-resource language context.
Attribution-guided query rewriting method that uses token-level explanations to improve neural retrieval system robustness.
Region-to-Image Distillation: technique to improve fine-grained visual perception in multimodal LLMs without repeated zooming operations.
RouterXBench: evaluation framework for routers in collaborative LLM systems that deploy smaller models locally and offload complex queries to cloud services.
Meta-cognitive architecture for agentic AI in cybersecurity enabling accountable autonomous decision-making under adversarial uncertainty.
Method to improve Direct Preference Optimization for LLM alignment by addressing issues when reference model prefers rejected responses.
Systematic evaluation of using LLMs to co-evolve domain-specific language definitions and instances during grammar changes.
Study of interaction patterns and role assignment between humans and LLMs in collaborative high-stakes decision-making systems.
Framework for dynamically selecting among LLMs in evolutionary agents to balance computational efficiency and reasoning capability.
Manifold-aware geometric approach for temporal domain generalization in LLMs under parameter-efficient constraints.
Agent guidance approach accelerates robotic reinforcement learning without human-in-the-loop bottlenecks.
Evaluation of context files (AGENTS.md) effectiveness for improving coding agent task completion on real repositories.
Analysis of differential privacy impact on federated spiking neural networks in distributed neuromorphic learning.
Study of class imbalance effects on deep learning vulnerability detection in source code.
Fourier-based generative pipeline for modeling crystalline materials using diffusion models with physical constraints.
ModelWisdom toolkit for TLA+ model visualization and automated repair of formal specification counterexamples.
Behavioral study examining user preferences for different LLM assistant modalities in multi-party negotiation scenarios.
DeepSight comprehensive toolkit for LLM and MLLM safety including evaluation, diagnosis, and alignment workflows.
Multi Graph Search algorithm for efficient motion planning in high-dimensional robotic systems.
Theoretical analysis of offline reinforcement learning under Q-approximation and partial coverage conditions.
Meta-Sel uses supervised meta-learning for efficient demonstration selection in in-context learning with computational budget constraints.
On-policy distillation with reward extrapolation extends KL-constrained RL to improve student model performance beyond teacher alignment.
Empirical study of AI coding agent adoption analyzing 2,901 PR contributions to open-source Android and iOS projects.
dVoting introduces fast voting mechanism for Diffusion LLMs enabling parallel token generation and efficient test-time scaling.
3DGSNav enhances vision-language models for zero-shot object navigation in embodied AI agents using 3D Gaussian splatting.
SAGEO Arena environment evaluates Search-Augmented Generative Engines and optimization techniques for web content visibility in AI responses.
VRB benchmark dataset for evaluating multimodal LLMs on visual reasoning tasks from primary education math problems.
DeepGen 1.0: lightweight 5B parameter unified multimodal model for competitive image generation and editing capabilities.
Studies whether neural world models internalize true physics or exploit shortcuts, proposing methods to avoid measurement corruption.
Distribution Discriminant Theory enables on-policy supervised fine-tuning for LLMs, bridging SFT-RL generalization gap.
Energy-aware continual learning method for spiking neural networks on neuromorphic vision systems with event-based cameras.
Olmix: framework for determining optimal data mixing ratios during language model training with practical design guidance.
ExtractBench: benchmark and methodology for evaluating LLM-based PDF-to-JSON structured data extraction at enterprise scale.
Technical curriculum on language-oriented AI literacy for translation and specialized communication industry professionals.
Studies implicit regularization in Langevin dynamics with projected noise for understanding stochastic gradient descent in symmetric models.
AttentionRetriever leverages LLM attention layers for long document retrieval, enabling better RAG without separate retrieval models.
UniT enables unified multimodal models to iteratively refine outputs through test-time scaling with chain-of-thought reasoning.
Vision-Language-Action model test-time verification reduces instruction-action misalignment more effectively than policy scaling.
SuperARC: AGI/ASI test based on recursive compression and mathematical randomness rather than human-centric questions.
SKATE: LLM evaluation framework where models generate and solve verifiable tasks for each other, treating evaluation as a competitive game.
Proposes design principles for explainable AI systems that integrate domain knowledge and contextual reasoning for better human comprehension.
Combines reinforcement learning with search-based path planners to optimize flight trajectory planning for emergency route recalculation.
Multi-agent system reasoning improvement through confidence-weighted consensus aggregation, inspired by democratic committee decision-making.
Method for improving LLM reasoning by optimizing training data logical complexity and entailment relationships rather than just factuality or diversity.
Human Behavior Atlas: Unified benchmark dataset for understanding psychological and social behaviors through multimodal AI, covering affective and cognitive states.
MARSHAL: Reinforcement learning framework for training LLMs in multi-agent competition and cooperation using self-play and strategic reasoning.
SMaRT: Strategy fusion framework combining multiple reasoning approaches for LLM task automation, maximizing performance and robustness over single-strategy methods.
SIEV framework: Evaluation method for LLM reasoning beyond correctness, modeling reasoning as dynamic dialectical trajectory rather than static chains.
DriveSafe: Domain-specific risk taxonomy and evaluation framework for LLM-based driving assistants addressing safety-critical automotive scenarios.
Theoretical analysis of multi-agent system performance under fixed inference budget; predicts regimes where agents help, saturate, or fail based on context windows and communication constraints.