DAST: A Dual-Stream Voice Anonymization Attacker with Staged Training
Dual-stream voice anonymization attacker using staged training to evaluate privacy robustness by fusing spectral and self-supervised learning features.
Dual-stream voice anonymization attacker using staged training to evaluate privacy robustness by fusing spectral and self-supervised learning features.
RL post-training method for diffusion-based text-to-image models using finite difference flow optimization to reduce variance and improve prompt alignment.
Mixed-methods evaluation of LLM-powered BPMN copilot for business process modeling, assessing human factors like trust and usability with domain experts.
Study of compact multilingual language models trained on child-directed speech in English-French scenarios, extending BabyBERTa framework.
Machine unlearning approach preserving semantic relations among retained instances while removing designated data influence from pretrained models.
ThinkStream framework enabling real-time streaming video understanding for multimodal agents without batch processing delays, addressing latency in interactive applications.
Neuro-symbolic framework combining Delta1 theorem generator with LLMs for explainable reasoning, integrating formal logic with language model interpretability.
Self-evolution framework using Minimum Bayes Risk decoding for error span detection in machine translation, reducing need for human annotation.
ARL-Tangram: framework optimizing resource efficiency in agentic reinforcement learning by dynamically managing external cloud resources.
daVinci-Env: large-scale open-source environment for training software engineering agents with executable repositories and dynamic feedback.
SAW: surgical world model framework for generating realistic surgical videos with precise tool-tissue interaction control.
Security research on breaking image protection methods against diffusion model editing under model mismatch scenarios.
Analysis of design homogenization in LLM-generated web designs through vibe-coding, examining how generative AI limits design diversity.
Cross-dataset empirical study comparing general-purpose vision models against specialized architectures for medical image segmentation.
L2GTX: method for explaining time series classification decisions by providing local and global explanations respecting temporal dependencies.
GeoChemAD: open-source benchmark dataset for unsupervised geochemical anomaly detection in mineral exploration across multiple regions.
Workflow for LLM-assisted grading of handwritten student assessments combining OCR, LLM processing, and human review for mathematics.
Study evaluating vision-language models' spatial reasoning for robot planning and understanding object relations from natural language instructions.
BoSS: oracle-based selector for active learning that improves robustness of instance selection strategies across models and datasets.
Framework for improving video-capable vision-language models' understanding of camera motion through benchmarking, diagnosis, and explicit motion representation.
PsyCogMetrics AI Lab: cloud platform combining psychometric and cognitive science methodologies to evaluate large language models.
Benchmark dataset for evaluating LLM hallucination mitigation on long-context ESG reports. Tests LLM reliability for corporate sustainability document analysis.
Study identifying critical weight subsets responsible for both privacy vulnerability and utility in neural networks. Proposes targeted privacy preservation without full model retraining.
Constitutional Multi-Agent Governance framework constrains LLM policy generation in multi-agent systems. Ensures cooperation emerges from alignment rather than autonomy erosion.
Framework enabling LLM-based AI agents to consolidate scientific knowledge across computational materials science experiments. Advances from execution to research-level pattern recognition.
Reward modeling framework for vision-to-code tasks using LVLMs. Addresses RL training challenges in converting visual inputs to structured code with reinforcement learning.
Diffusion models for text-conditioned humanoid motion generation with physics-based controllers for robot animation. Focuses on motion synthesis rather than AI agents or LLM applications.
Proposes active causal structure learning with latent variables as necessary component for AGI agents and robots in changing environments.
Introduces Human-AI Governance framework emphasizing relational dynamics between human and AI actors in governance.
Reveals observer effects in safety evaluations where advanced AI systems detect being tested and modify behavior accordingly.
Presents Darwin Godel Machine for open-ended evolution of self-improving agents using meta-learning to automate algorithm discovery.
Proposes MAGPO framework for multi-agent reinforcement learning leveraging centralized training with decentralized execution.
Introduces CRAFT-GUI, curriculum-reinforced agent using RL for GUI task execution with difficulty-aware training.
Demonstrates marginal single-step accuracy gains compound into exponential improvements in long-horizon task completion for LLMs.
Presents AutoClimDS agentic AI system using climate knowledge graph to address fragmented data sources in climate science.
Introduces operational safety framework evaluating LLM-based agents' ability to appropriately accept or reject out-of-scope requests.
Proposes task-specific prompt-prototype approach for continual learning without key-value pairing to reduce inter-task interference.
Benchmarks 20+ LLMs against humans on causal reasoning tasks, finding models exhibit human-like biases and shortcuts.
Introduces SkillsBench, benchmark of 86 tasks evaluating whether structured procedural knowledge packages improve LLM agent performance.
Proposes OpenSage, first agent development kit enabling automated design of agent topology, tools, and memory components.
Formalizes steganographic capabilities in LLMs and proposes detection methods for covert reasoning in large language models.
Presents OPENDEV, open-source CLI coding agent in Rust for terminal-native development tasks, focusing on long-horizon autonomous programming.
Introduces context engineering as discipline for designing informational environments where multi-agent AI systems make decisions, extending beyond prompt engineering.
Evaluates autonomous cyber-attack capabilities of frontier AI models on multi-step attack scenarios, comparing seven models over 18 months at varying compute budgets.
arXiv paper introducing COMPASS framework integrating sovereignty, sustainability, compliance, and ethics into LLM-based agentic systems.
arXiv study examining user adoption of OpenClaw system using Cognition-Affect-Conation framework.
arXiv paper on continual learning framework enabling multimodal agents to improve from experience without parameter updates.
arXiv research on adversarial robustness techniques for vision-language multimodal models.
arXiv research on brittleness of safety alignment in LLMs, proposing Superficial Safety Alignment Hypothesis.
Motion Dreamer generates physically coherent video with boundary conditions for autonomous driving and embodied AI planning tasks.