NeuroCogMap Reveals Cognitive Organization of Large Language Models
arXiv research introducing NeuroCogMap to reveal cognitive organization within large language models through functional mapping.
arXiv research introducing NeuroCogMap to reveal cognitive organization within large language models through functional mapping.
arXiv research showing safety alignment metrics in text-to-image diffusion models may misrepresent utility due to coarse global metrics.
arXiv paper on orbital ray-conditioned 3D foundation models for satellite imagery reconstruction using multi-view data.
RL framework for quadruped gait learning using Signal Temporal Logic specifications instead of hand-crafted rewards.
Iterative video retrieval and reasoning system with soft query refinement for temporal grounding and inter-video analysis.
LLM-based hard negative sampling technique for training two-tower retrieval models in large-scale recommendation systems.
Analysis method for critic model complexity in actor-critic RL using spectral effective-rank entropy.
Physics-informed deep learning framework for spatiotemporal dynamics prediction with multi-resolution analysis.
Security analysis of function-calling LLMs showing jailbreak vulnerabilities through simulated moderation traces.
Active learning framework for personalizing diffusion models via user preference feedback in recommendation systems.
Benchmark for testing spatial counterfactual reasoning in vision-language models using smartphone photo triplets.
Native Metal inference runtime for LLMs on Apple Silicon with optimized kernel fusion and memory management.
Self-supervised learning method for 4D point cloud representations using cross-modal distillation from 2D models.
LLM fine-tuning approach combining imitation learning and RL for multi-step molecular optimization reasoning.
RL method for training few-step diffusion models via policy optimization with stochastic flow-maps.
Study evaluating cross-domain generalization of lightweight ML models for IIoT intrusion detection systems.
Novel deep learning architecture combining group-equivariance with hyperbolic geometry for visual representation learning.
Empirical study of AI design patterns in open-source software repositories to validate their real-world prevalence and utility.
Framework for auditing whether deleted facts persist in limited-memory LLMs through residual parameters or retrieval artifacts.
Method for discovering novel object categories in open-world settings by identifying latent concepts in vision model representations.
ML research on smooth loss transitions during neural network adaptation under distribution shift, improving representation preservation during fine-tuning and RL.
Lightweight backbone-agnostic mask adapter for fair comparison of transformer backbones in image segmentation tasks.
LLVM-Bench: benchmark with 423 real-world issues for evaluating LLMs on compiler problem resolution and issue triage.
Self-conditioning technique for flow-based language models using fixed-point flows to enhance text generation in few-step generators.
LLMs guide ODE discovery and parameter inference from small-cohort aggregate data for mechanistic clinical modeling in rare diseases.
Measures hallucinated citations in peer-reviewed papers generated by LLMs, demonstrating how unsupported references survive conference review.
Vision-language pretraining without contrastive objectives, adopting non-contrastive methods for dense prediction and frozen visual backbones.
Formative study of operationalization failures when LLMs generate queries and construct analytical workflows, identifying semantic gaps in data systems.
Studies gradient-based inversion techniques to recover input token sequences from decoder-only LLM hidden states through embedding-space optimization.
Adaptive token length reduction for large reasoning models based on confidence estimation to improve inference efficiency.
Research paper on MAGNET: Multi-agent framework for long-form narrative generation using persona-grounded character agents.
Research paper on Generative Meta-Learning with Human Feedback (GMHF). Framework for domain adaptation using expert guidance.
Post-training pruning method for Diffusion Transformers addressing computational overhead in image generation models.
Learning cardiac motion priors for implicit neural representations to improve motion field estimation and optimization speed.
LLM-powered agent simulation for semantic trajectory analysis in zoned environments, modeling human movement with behavioral patterns.
SWE-Doctor guides LLM-based software engineering agents using runtime diagnosis from bug reproduction tests for patch generation.
Logit-contribution scoring method identifies attention heads in LLMs that synthesize non-literal answers from context in long-context settings.
Graph-based training-free framework for inferring reading order in complex historical document layouts with interleaved text streams.
Framework for LLM-based conversational agents to dynamically adapt personality and persona based on interaction context.
DART-VLN training-free test-time control framework for discrete vision-language navigation addressing memory staleness and loop avoidance.
MemSyco-Bench benchmarks sycophancy in LLM agent memory systems where retrieved memories cause over-alignment with users.
Analyzes staleness effects in asynchronous RLHF training with stale rollouts, establishes scaling laws for learning rate adjustment.
LongVQUBench benchmark evaluates vision-language models on long-term video quality assessment with temporal degradation patterns.
Study of AI-mediated software engineering focusing on inspection, correction, and maintainability when agents generate code at scale.
CausalMix method treats data mixture optimization as causal inference for LLM training, addressing adaptive data distribution shifts.
FAR framework enables robots to learn from test-time failures and adapt behavior autonomously without human intervention.
Multimodal university chatbot using retrieval-augmented generation for domain-specific queries and institutional information access.
Mechanistic interpretation of Muon optimizer for neural network training as implicit residual connection.
Autonomous scientific discovery system using iterative meta-reflection for hypothesis generation and validation without predefined research constraints.
Framework for managing dependencies and provenance in LLM agent skill supply chains to prevent duplicated dependencies.