PolicySim uses LLM-based agents to simulate social platform dynamics and evaluate intervention policies before deployment, addressing echo chambers and polarization.
Proves KV cache in transformers is redundant; keys and values can be deterministically recomputed from residual streams with zero error, enabling efficient inference.
GoAgent proposes a method to automatically generate optimal communication topologies for LLM-based multi-agent systems, enabling task-specific agent group structures for complex problem-solving.
AIGQ end-to-end generative framework for e-commerce query recommendation addressing cold-start and shallow semantics issues.
MOSS-TTSD system for text-to-spoken dialogue generation handling turn-taking, acoustic consistency, and long-form stability.
Uncertainty-aware prototype learning with variational inference for few-shot 3D semantic segmentation of point clouds.
HOP3D framework for few-shot 3D point cloud segmentation using hierarchical orthogonal prototypes to address stability-plasticity trade-off.
Fine-tuning framework (SeGroS) for unified multimodal models that resolves granularity mismatch through semantically-grounded supervision.
Semantic role classification approach using analogies over frame elements and lexical units in FrameNet datasets.
Statistical metric using semantic category distribution to distinguish human-written from LLM-generated dialogue via lightweight lexical analysis.
Proposes selective-complementary reinforcement learning for test-time LLM enhancement, addressing weak consensus vulnerability in majority voting pseudo-rewards.
Integrates meta-features with knowledge graph embeddings for pipeline performance estimation and dataset similarity estimation in meta-learning.
Analyzes domain-spatiality patterns in fitness landscapes for configuration tuning, combining static and dynamic analysis for better explainability.
ATCG module uses analogical reasoning from labeled knowledge for generalized category discovery on fine-grained visual classification.
Proposes span-level evaluation methodology for machine translation meta-evaluation, extending beyond scalar scores to error detection and severity assessment.
RAM recovers 3D human motion from video using motion-aware semantic tracking, adaptive Kalman filtering, and temporal HMR module.
HiPath framework aligns vision-language models for structured pathology report prediction using frozen UNI2 and Qwen3 backbones.
Demonstrates injection attack on OpenClaw autonomous coding agents through bootstrapped guidance in lifecycle hooks, revealing security vulnerabilities.
X-World simulator generates realistic driving scenarios for scalable evaluation of vision-language-action autonomous driving policies without real-world testing.
Addresses LLM post-training capability ceiling by reintroducing Markov states to reinforcement learning, enabling discovery beyond pre-trained pattern refinement.
Derives long-range Coulomb corrections for machine-learning Hamiltonians in electronic structure prediction, improving accuracy over density-functional theory.
Proposes Detached Skip-Links and R-Probe to decouple feature aggregation from gradient propagation, improving multimodal LLM OCR task performance.
Longitudinal field study of Chiron platform coordinating human-AI agents across software modernization delivery stages: analysis, planning, implementation, validation.
CoverageBench evaluates information coverage of ad-hoc retrieval algorithms in RAG systems, proposing metrics beyond precision and recall.
LoASR-Bench benchmark evaluates speech language models on automatic speech recognition across low-resource languages. Addresses gap in LLM ASR performance understanding for underrepresented language families.
Reinforcement learning fine-tuning methods for financial time series forecasters with supervised-to-RL backpropagation and transfer learning benefits.
llvm-autofix agentic harness enables LLM agents to autonomously fix compiler bugs with compiler-specific tools and cross-domain expertise.
LLM-based semantic data integration for aerospace electronic component qualification information retrieval across departmental data silos.
Empirical comparison of SFT, DPO, and staged training on GPT-2-scale models for paraphrase detection and sonnet generation tasks.
Temporal abstraction analysis addresses spectral mismatch in forward-backward representations for continuous space successor representation learning.
Lambda-RLM framework solves long-context LLM bottleneck using lambda-calculus for verifiable recursive subproblem decomposition.
Var-JEPA bridges predictive and generative self-supervised learning by showing JEPA design mirrors probabilistic modeling structure.
Adapt4Me web application uses Bayesian active learning for personalized ASR adaptation to non-normative speech without expert supervision.
Chain-of-Adaptation framework for domain-specific vision-language model fine-tuning preserving pretrained reasoning capabilities.
Automated multi-objective long-tail attacks on LLMs exploiting low-resource languages and encrypted data vulnerabilities.
Six-agent AI system for cybersecurity risk assessment following NIST CSF framework, automates profiling, asset mapping, threat analysis, and risk scoring.
Attention-based pooling for HAL semantic representations improves text classification over mean pooling.
Design-OS specification-driven framework for engineering system design with AI assistance, including control-systems case study.
Semantic token clustering method for efficient uncertainty quantification in LLMs without repeated sampling or auxiliary models.
CRISP framework uses vision-language models as social critics to enable autonomous robot self-refinement of social behaviors.
Demonstrates that chain-of-thought faithfulness measurements vary significantly by classifier choice, challenging objective faithfulness claims.
LLM-based AI agents autonomously execute high energy physics analysis pipelines including event selection, background estimation, and statistical inference.
Adaptive frame selection method for long-video understanding in vision-language models to reduce computational bottlenecks.
Multi-modal contrastive learning to improve ML generalization in cybersecurity tasks and reduce shortcut learning.
VideoSeek: long-horizon video agent using tool-guided seeking and video logic flow. Active frame selection reduces computation vs dense frame sampling.
LumosX framework for personalized video generation with identity-attribute alignment. Diffusion-based approach for fine-grained face-attribute control.
Taxonomy and benchmark for VLM image tampering detection. Pixel-grounded evaluation of vision-language models on edit detection tasks.
Hard preference sampling for LLM alignment with human preferences. Improves Plackett-Luce and Bradley-Terry models for harmful content handling.
Average reward RL for omega-regular and mean-payoff objectives. Automated compilation of formal behavioral requirements into learning objectives.
Deep RL for multi-objective combinatorial optimization with conditional computation. Uses weight vectors to explore solution space efficiently.