CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
Training-free token pruning technique for efficient vision-language models optimized for pixel grounding tasks.
Training-free token pruning technique for efficient vision-language models optimized for pixel grounding tasks.
Few-shot action recognition using semantic-temporal adaptive learning with vision-language models.
Video understanding agent with recursive tool-use, meta-augmented grounding, and fine-grained operations for temporal reasoning.
LLM distillation method using teacher guidance and Reverse KL to improve student model learning from divergent distributions.
Information-theoretic framework for visual evidence selection in multimodal RAG systems based on utility to downstream reasoning.
Study analyzing readability of LLM-generated code compared to human-written code and effects of prompt design.
Exploration algorithm balancing uncertainty reduction with expected improvement in large action spaces.
Parallel multi-turn medical dialogue dataset spanning English and nine Indic languages with LLM-generated synthetic conversations.
Credit attribution framework for optimizing multi-agent LLM systems through contrastive learning to improve agent configuration.
Analysis of embodied AI agents' limitations in vision-language navigation when transitioning from simulation to real-world deployment.
Interpretability research tracing how LLMs represent behavioral traits like sycophancy as linear directions in internal activations.
Genetic algorithm-based DoS attack exploiting LLM reasoning model vulnerabilities by inducing excessive inference through incomplete inputs.
Mechanistic study of how LLMs implement persona-dependent preferences using linear probing of model internals.
Runtime substrate architecture for foundation-model software agents mediating agent-environment interaction and code generation reliability.
Test-time self-training approach enabling LLM parameter updates at inference for query-specific adaptation.
Token pruning method for vision-language models using group-relative importance ranking for computational efficiency.
Evaluation of off-the-shelf LLMs for legal document annotation on Danish asylum credibility assessment task.
Multilingual foundation model framework for detecting reclaimed slurs in social media across three languages.
Reinforcement learning approach using flow-based models as policies with stable gradient-based optimization methods.
Bimanual robot control framework balancing independent arm perception with coordinated interaction through visuomotor learning.
Method for discovering localized model calibration failures beyond global reliability metrics using structured analysis.
Study of chain-of-thought in-context learning with many-shot examples showing scaling behavior comparable to fine-tuning on reasoning tasks.
LLM-based high-level synthesis code generation using comparative reward reinforcement learning for hardware optimization.
Inference-time alignment method for LLMs using temperature adjustment to mitigate reward hacking in model outputs.
On-device PII redaction pipeline using small language models with few-shot prompting for privacy-preserving text substitution.
Robotic foundation model training method addressing temporal heterogeneity in action sequences through weighted optimization.
arXiv paper on AttenA+: improving robotic foundation models by addressing temporal heterogeneity in manipulation tasks.
OpenAaaS: open-source agent-as-a-service framework for distributed materials-informatics research using LLMs and autonomous agents.
arXiv paper on NAACA: training-free audio language model architecture with oscillatory working memory for salience-driven attention.
arXiv paper analyzing whether low-rank pre-training methods for LLMs generalize comparably to full-rank training.
Synthetic hierarchical language model with provable scaling laws and benefits of multi-step reasoning.
RTLC three-stage prompting improves LLM-as-judge accuracy for evaluation without fine-tuning or external tools.
Method using canary tokens to detect and identify web scraping by AI systems for LLM training data collection.
Fine-tuned compact LLMs generate children's reading stories with controllable difficulty and safety constraints.
Critical analysis of 'human in the loop' as safety mechanism for AI systems, examining its limitations and misuse.
KVServe compresses KV cache dynamically for disaggregated LLM serving to reduce communication bottlenecks.
Quantization techniques for weight-only post-training of LLMs using available covariance matrix information.
Method for detecting step-level hallucinations in LLM reasoning by analyzing hidden state geometry during inference.
Study evaluating whether LLMs consistently understand semantic content in High-Level Message Sequence Charts for software design.
MinT infrastructure system for efficient LoRA fine-tuning and serving of millions of LLM variants using shared base models.
LMPath uses language models to generate semantic-aware exploration paths for autonomous UAV search missions.
Research on improving LLM evaluation reproducibility by modeling annotator biases in human rating systems for AI safety assessment.
Neurosymbolic approach combining LLMs with SMT solvers to audit natural-language software requirements for ambiguity and inconsistency.
Negation Neglect: LLMs fail to learn negations during finetuning, believing falsified claims despite recognizing them as false in context.
EVA-Bench: Evaluation framework for voice agents addressing realistic conversation simulation and voice-specific failure mode measurement.
WARDEN: Language model for transcribing and translating endangered Wardaman language using only 6 hours of annotated audio data.
ChatSR: Multimodal LLM for scientific formula discovery from graphs and images with specialized scientific data understanding.
Taxonomy and survey of AI safety landscape for large language models covering design, development, adoption and deployment impacts.
Language model networks study: pre-trained LMs as reusable nodes in inference systems with learned communication patterns instead of natural language.
Block-wise adaptive caching technique to accelerate Diffusion Policy inference for real-time robotic control.