LEVI: Stronger Search Architectures Can Substitute for Larger LLMs in Evolutionary Search
LEVI framework showing stronger search architectures can substitute for larger LLMs in evolutionary search, reducing computational costs.
LEVI framework showing stronger search architectures can substitute for larger LLMs in evolutionary search, reducing computational costs.
Sparse autoencoder feature steering amplifying Dark Triad personality traits in Llama-3.3-70B to study antisocial behavior circuits.
EvoPref multi-objective evolutionary algorithm optimizing LoRA adapters for LLM alignment across helpfulness, harmlessness, and honesty objectives.
QD-LLM framework using parameter-efficient neuroevolution to evolve prompt embeddings for diverse LLM generation via Quality-Diversity optimization.
CrossVL method for vision-language detection across different viewpoints using complexity-aware feature routing and curriculum learning.
Insight Android accessibility service using LLMs for natural language interaction and real-time screen summarization for blind and visually impaired users.
LEAD framework for reducing verbosity in reasoning models' Chain-of-Thought trajectories through length-efficient adaptive reasoning.
Oracle Poisoning attack class where adversaries corrupt knowledge graphs queried by AI agents at runtime via tool-use, causing incorrect reasoning.
CalBench evaluation environment for multi-agent LLM coordination through calendar scheduling with privacy constraints.
Controlled study of MXFP4 quantization in transformer pretraining, analyzing stability across forward propagation, activation gradients, and weight gradients.
Fine-tuned Florence-2 vision-language model using LoRA for extracting structured fashion attributes from clothing images as JSON.
Score-based conditional energy model for inference in hybrid Bayesian networks with discrete and continuous variables.
Study of routing traces in Attention-Residual transformers for post-hoc calibration, examining whether routing signals improve uncertainty estimates.
Unified geometric framework using flag varieties to theoretically explain alignment phenomena in deep networks including gradient flow and neural collapse.
UFO framework for robust continual graph learning addressing noisy annotations in evolving graphs through flow-oriented methods.
Nautilus Compass: black-box persona drift detector for production LLM coding agents that tracks constraint violations and memory degradation over long sessions.
Method enabling efficient online learning in transformers through continuous latent contexts for multi-turn adaptation and feedback-driven decision making.
SVAR-FM framework for time series causal discovery using physics-based simulators to generate interventional data with flow matching.
EgoMemReason benchmark for evaluating memory and reasoning in ultra-long egocentric video understanding for embodied agents and smart glasses.
Key-Value Means (KVM): O(N) attention mechanism supporting fixed or growing state for long-context transformers with subquadratic prefill performance.
Identifies 'Cartesian Shortcut' vulnerability in vision reasoning benchmarks where MLLMs exploit grid-based layouts through explicit coordinate discretization.
Theoretical analysis explaining why sparse autoencoders (SAEs) show layerwise scaling variations through manifold geometry, advancing understanding of activation space structure.
VALDI benchmark demonstrating 'Pseudo-Deliberation' failure mode where LLMs show reasoning without behavioral alignment, revealing value-action gaps.
Position paper identifying 'Agentic Denominator Gaming' threat where malicious actors deploy AI agents to flood academic conferences with low-quality papers exploiting stable acceptance rates.
NaiAD dataset of 58,999 ad-embedded LLM responses with evaluation metrics for studying LLM-native advertising balancing user experience and platform revenue.
VIGOR: reward model for LLM reinforcement learning that eliminates need for external verifiers by using intrinsic gradient-norm signals, improving scalability to new tasks.
Self-training method using team-based self-play with dual adaptive weighting to improve LLM alignment while reducing dependency on human-labeled data and addressing synthetic data quality issues.
Inference-time pruning method reducing unnecessary tool calls in LLM reasoning systems while maintaining performance.
GPU-accelerated Boruta feature selection algorithms for high-dimensional datasets with improved computational efficiency.
Self-improvement framework for LLMs using intrinsic rewards without verifiers, enabling autonomous evolution on open-ended tasks.
Small 0.3B multilingual NER model for detecting 42 PII entity types across languages and document formats.
Analysis of attention drift phenomenon in speculative decoding drafters showing attention shifts from prompt to generated tokens over speculation chains.
Continual Harness framework for online adaptation of embodied agents with iterative refinement, demonstrated on Pokemon gameplay.
Specification inference tool for Move Prover combining weakest-precondition analysis with agentic Claude coding for reduced boilerplate.
Tag-based few-shot example selection method for prompting LLMs to generate causal factors and preventive measures from medical incident reports.
LLM-based framework for automated psychological crisis assessment from speech for mental health hotline support.
Vision for unified memory paradigm for agentic AI in 6G radio access networks to bridge semantic bottleneck in disaggregated architectures.
C-BPO framework for personalizing LLMs using binary preference feedback with inter-user calibration.
Guidance mechanism for stochastic interpolant robot policies enabling test-time steering without retraining for dynamic objectives.
Portable specification for multi-agent coordination as self-improving skills distributed across agent frameworks.
NCO plugin for constraining LLM decoding to prevent undesirable outputs like profanity and PII during generation without post-processing.
Metis framework reformulates LLM red teaming as policy optimization using adversarial MDPs for improved jailbreak discovery.
Test-time adaptation framework for vision-language-action robotic models that uses successful execution history to improve closed-loop reliability.
ViSRA framework for probing spatial reasoning in multi-modal LLMs through video-based agent without requiring model retraining.
Adaptive perception system for autonomous driving that dynamically allocates computation based on scene complexity.
MicroWorld framework enhances MLLMs for microscopy via multimodal attribute graphs to bridge domain gap with limited training data.
Study investigating whether scaling vision models improves localization-based explanation quality across ResNet, DenseNet, ViT architectures.
NLP method for detecting and analyzing contradictions in scientific peer reviews using fine-grained analysis beyond sentence pairs.
Security framework addressing SQL injection vulnerabilities in LLM-driven database interfaces through prompt-to-SQL translation.
MTA-RL framework combines transformers and reinforcement learning for autonomous driving with 3D scene understanding.