MURPHY: Multi-Turn GRPO for Self Correcting Code Generation
MURPHY: Multi-turn reinforcement learning framework for self-correcting code generation combining group relative policy optimization with execution verification.
MURPHY: Multi-turn reinforcement learning framework for self-correcting code generation combining group relative policy optimization with execution verification.
EGGROLL: Scalable evolution strategies algorithm using low-rank approximations to improve training efficiency of black-box optimization on GPUs.
ESPO: Entropy importance sampling policy optimization for stable and efficient token-level RL training of LLMs on complex reasoning tasks at scale.
RRPO: Robust reward policy optimization framework preventing reward hacking in LLM-based emotional text-to-speech by addressing vulnerability of vanilla reward models.
ArtistMus: Benchmark dataset for retrieval-augmented music question answering grounded in artist metadata to evaluate LLMs on music-related reasoning tasks.
IntentMiner: Privacy attack exploiting Model Context Protocol servers to extract user intents from LLM tool calls, revealing new security vulnerabilities in agentic AI systems.
IPNet: Neuromorphic architecture using magnetic tunnel junction intrinsic plasticity to implement human-like working memory with reduced energy costs.
Agent-based model simulating adaptive firm behavior in spatial double-auction markets to understand emergence of industrial symbiosis under socio-spatial constraints.
Multi-LLM validation framework for thematic analysis combining Cohen's Kappa and semantic similarity metrics to improve reliability of LLM-based qualitative research coding.
Nightjar: Dynamic adaptive speculative decoding method that adjusts verification overhead based on request load to optimize LLM inference throughput and latency.
Retrieval-augmented generation approach addressing domain shift in low-resource neural machine translation using context volume from limited corpora.
Unified framework for LLM alignment decoupling sampling and optimization geometry across PPO, DPO, IPO algorithms and variants.
LLM-based AI agents with hypothesis-driven cognition for improved software bug localization by analyzing code component relationships.
Agentic operationalization of DISARM framework for investigating foreign information manipulation on social media across NATO allied partners.
Machine unlearning method for Mixture-of-Experts LLMs using geometric router constraints to erase knowledge rather than redirect queries.
Framework for interpreting and explaining emergent extreme events in LLM-powered multi-agent systems to improve safety and transparency.
Self-distillation approach for reinforcement learning leveraging rich textual feedback from verifiable environments to improve credit assignment in code/math tasks.
Benchmark for open-domain video shot retrieval using LLMs for understanding editing requirements and retrieving keyframe-oriented shots.
Analyzing semantic geometry in LLM hidden states versus behavioral similarity through psycholinguistic experiments across eight instruction-tuned models.
Unified retrieval-augmented generation framework for query auto-completion combining ranking and generation to reduce hallucination and improve coverage.
Reinforcement learning technique filtering irrelevant tokens to improve LLM policy optimization by focusing on contextually relevant action spaces.
Case studies of Google's Gemini models assisting scientific research including mathematical discovery and routine task automation.
Studying knowledge distillation from privileged information in language models for multi-turn agentic environments, addressing inference-time capability transfer.
Demonstrates that steering vectors in LLMs are fundamentally non-identifiable due to large equivalence classes of behaviorally identical vectors.
Steering-based jailbreak method against aligned LLMs requiring less computation than white-box approaches but maintaining stealth.
Benchmark for evaluating LLM performance on financial analysis and tracking using SEC filings with multi-document synthesis.
Analyzes errors and limitations in Code World Models that simulate program execution by predicting runtime state.
Studies language understanding through paraphrase generation and detection capabilities in language models.
Domain-specific language for predictive modeling on relational databases covering missing values and future predictions.
Derives scaling laws for massive-scale recommendation systems through unified architecture design and efficiency improvements.
Privacy protection method for mobile GUI agents using anonymization to mitigate exposure of sensitive data during screen processing.
Causal continuous rotary positional encoding improving 3D feature alignment in multimodal LLMs for spatial reasoning tasks.
Language-Action pre-training framework enabling zero-shot transfer of robot policies across different embodiments without fine-tuning.
Data attribution method to trace and mitigate undesirable emergent behaviors in LLM post-training by identifying responsible datapoints.
3D Gaussian Splatting simulation framework for training visual navigation policies in dynamic environments with moving obstacles.
Method to improve fine-grained visual perception in multimodal LLMs by distilling zoomed regions without repeated inference calls.
Benchmark and methodology for evaluating LLM-based PDF-to-JSON extraction at enterprise scale with complex schema requirements.
Formalizes agent skills as composable packages enabling LLMs to extend capabilities dynamically without retraining, covering architecture, acquisition, and security.
MedXIAOHE medical vision-language foundation model with entity-aware continual pretraining achieving state-of-the-art on medical benchmarks.
Prior-guided symbolic regression framework incorporating scientific constraints to discover interpretable equations consistent with physical principles.
Directional Concentration Uncertainty framework for flexible uncertainty quantification in generative models across tasks and modalities.
Blueprint system for multimodal retrieval in engineering archives using layout-aware VLM OCR and identifier normalization.
Comparative study of ML/DL architectures on MNIST-1D dataset for distinguishing neural network performance on controlled benchmarks.
Review and proposed metrics for evaluating multi-iteration active learning query methods and sample selection strategies.
Research showing language model circuits are prompt-specific within tasks, not uniform across prompts, using causal communication analysis.
Federated learning framework for nonlinear temporal dynamics using graph attention for interpretable cross-client relationships without sharing raw data.
Research identifies rank collapse phenomenon in federated low-rank adaptation with heterogeneous clients, proposing mitigation for privacy-preserving fine-tuning.
TrasMuon optimizer enhances Muon-style methods with trust-region adaptive scaling to address training sensitivity and high-energy bursts.
Novel optimization framework for monotone non-convex functions unifying DR-submodular and OSS functions with theoretical convergence analysis.
Study demonstrating singular vectors of attention heads align with learned features in language models, providing mechanistic interpretability foundations.