The Ideation Bottleneck: Decomposing the Quality Gap Between AI-Generated and Human Economics Research
Decomposes quality gap in AI-generated economics papers into research idea quality and execution quality using fine-tuned LMs.
Decomposes quality gap in AI-generated economics papers into research idea quality and execution quality using fine-tuned LMs.
Method for learning compositional latent actions from video for embodied AI using structural priors on physical motion.
Multi-stage pipeline integrating experimental design with ML surrogates for systematic exploration of agent-based models.
Benchmark for evaluating creative problem-solving in LLMs combining logical reasoning, lateral thinking, and commonsense knowledge.
Multi-modal agentic systems for iterative image editing; addresses quality degradation in multi-turn edits through image replication methods.
Privacy-preserving classroom attention analysis pipeline using OpenPose and gaze estimation. LLM evaluates student engagement from skeletal data without storing video.
Connects generative AI (diffusion, SDEs) to computational mechanics for material design. Reviews mechanisms and interpretability of generative models.
Weight-space arithmetic enables zero-shot quantization robustness transfer between models without training. Novel ML technique for model optimization.
AEGIS: Scaling homomorphic encrypted transformer inference via hybrid parallelism on multi-GPU. Privacy-preserving ML optimization, niche application.
MetaSAEs: Introduces decomposability penalty for training sparse autoencoders with atomic latents. Improves alignment and safety-relevant applications.
Compares RAG vs standard approaches for Agile story point estimation in sprint planning. arXiv study on LLM application.
TRACE: Study on how LLMs allocate trust between conflicting code, documentation, and tests. Evaluates trustworthiness in AI-assisted software engineering.
Introduces vocabulary dropout technique to solve diversity collapse in co-evolutionary LLM self-play curriculum learning. arXiv paper with novel method.
LLM-powered evolutionary search automatically discovers unsupervised uncertainty quantification methods as Python programs for claim verification.
Fine-tuning approach adapting DeepSeek-OCR-2 for optical chemical structure recognition by formulating task as image-to-text.
Study of brain-LLM alignment during creative divergent thinking tasks, measuring correlation between model performance and human neural activity.
VisionClaw wearable AI agent on Meta Ray-Ban glasses combining egocentric perception with speech-driven task execution via OpenClaw agents.
Sim2Real-AD framework for zero-shot sim-to-real transfer of VLM-guided RL policies from CARLA simulation to physical autonomous vehicles.
Dynamic model analyzing productivity-skill tradeoffs when workers use AI tools, decomposing productivity effects into expertise-dependent and independent channels.
Taxonomy of LLM-based coding agent architectures analyzing scaffolding code patterns including control loops, tool definitions, and context strategies.
LangFIR uses sparse autoencoders on monolingual data to discover language-specific features for steering LLM output language without parallel corpora.
AgenticFlict dataset of merge conflicts from AI coding agent pull requests on GitHub, studying integration challenges in collaborative AI-assisted development.
Video diffusion framework (CRAFT) for generating synthetic bimanual robot manipulation demonstrations with temporal coherence.
Phase-aware suppression method to reduce hallucinations in Vision-Language Models without iterative optimization overhead.
SecPI framework for secure code generation using reasoning LLMs through security reasoning internalization, addressing inference-time vulnerability mitigation.
Neural method for black-box global optimization using iterative refinement from noisy samples, addressing multi-modal function optimization.
LLM-based approach for multi-file repository code generation with executable validation, addressing dependency resolution and integration challenges.
LiveCoder framework for repository-level code generation preserving and reusing task-specific state across multiple LLM attempts.
Generative foundation model for multimodal histopathology that imputes missing modalities from incomplete medical data.
Method for stable unsupervised self-evolution of multimodal LLMs using continuous softened retracing resampling for feedback quality.
Microservice system using NLP and deep learning to automate classification of citizen appeals in government services.
Unlocks prompt infilling in masked diffusion language models by applying full-sequence masking during supervised finetuning.
LightThinker++ enables LLMs to dynamically compress intermediate reasoning thoughts into compact representations for efficiency.
Uses LLMs to capture semantic relationships for tail-item sequential recommendation, addressing sparse interaction problem.
CREBench evaluates LLMs on cryptographic binary reverse engineering, assessing capabilities for vulnerability discovery and malware analysis.
Research identifying limitations in universality of linear truth directions in LLM activation spaces across different settings.
Study measuring human ability to distinguish LLM-generated news from human-written content across six LLM models.
AutoReSpec uses LLMs to generate formal specifications for programs, addressing syntax and logic errors through techniques for complex control flow.
Neuro-symbolic framework for robot manipulation using vision-language models and autonomous domain construction.
Method for discovering repeated attention patterns in large language models at scale for mechanistic interpretability.
Automated framework for research-level mathematical problem solving combining LLMs with formal verification to reliably resolve conjectures and verify proofs.
Representational collapse in multi-agent LLM committees: measurement of similarity showing agents produce redundant rationales despite different role prompts, with diversity-aware consensus.
k-Maximum Inner Product Attention for efficient graph transformers, reducing quadratic complexity while maintaining expressiveness for large-scale graphs.
Analysis of analogical reasoning in LLMs comparing probed representations with prompted performance, revealing limitations in latent abstraction and generalization.
Field experiment on LLM agent providing iterative personalized behavioral nudges for electricity and hot-water conservation across intervention rounds.
I-CALM: prompt-only intervention reducing LLM hallucinations by incentivizing confidence-aware abstention through reward scheme announcements and humility principles.
DC-Ada: reward-only decentralized adaptation for heterogeneous multi-robot teams, adapting frozen policies to mismatched sensor configurations.
Secure-by-design GenAI framework for cloud security and forensics using LLMs with defenses against prompt injection and forensic rigor requirements.
Spatio-temporal sparse autoencoders for interpretable video representation learning, using contrastive objectives and hierarchical grouping to preserve temporal coherence.
Multi-turn decision making framework for goal-oriented conversational systems balancing information acquisition and target commitment under user intent uncertainty.