TimEE: End-to-end Time Series Classification via In-Context Learning
In-context learning approach for time series classification that eliminates separate feature encoder training and enables label exploitation at inference.
In-context learning approach for time series classification that eliminates separate feature encoder training and enables label exploitation at inference.
Study of how hallucinated content propagates through reasoning stages in vision-language models and affects downstream inference.
Asynchronous reinforcement learning system for LLM post-training optimized for long-horizon agentic tasks with improved training stability.
Interactive AI agent framework for structural design that explores alternatives and refines solutions while satisfying spatial, mechanical, and cost constraints.
Federated learning approach using collaborative synthetic data generation for knowledge transfer across distributed clients with divergent data distributions.
Multi-component simulator for autonomous driving that synthesizes corner cases combining visual representation, scene reasoning, and vehicle control.
Systematic review of governance challenges and frameworks for agentic AI systems capable of autonomous planning and task execution.
Method for improving confidence estimation in LLMs by tracking confidence evolution during generation for better deployment in tool use and adaptive systems.
DiaLLM: Method for improving dialectal English generation in open-weight LLMs through continual pretraining and alignment.
Sample-efficient RLHF method for diffusion models using selective timestep weighting and advantage-based replay.
Jailbreak: Agentic approach to bypass database engines by directly reading storage files for high-performance columnar analytics.
Continuous-query limited memory language models that externalize factual knowledge to knowledge bases during pretraining and generation.
Framework for mechanistically explaining structure-property relationships using deep learning with physical constraints and scientific principles.
Method for domain-independent planning that uses LLMs to automatically synthesize heuristics from problem definitions.
Deep reinforcement learning approach for universal robot control using modular recurrence and contextual MDPs across different morphologies.
Study using satellite imagery and LLM-generated text descriptions to investigate socioeconomic indicators in poverty mapping.
LiveOIBench: Large-scale benchmark of competitive programming problems for evaluating LLM coding capabilities with comprehensive test coverage.
AGAPI-Agents: Open-source agentic AI platform integrating LLMs with 28 scientific tools for accelerated materials design.
Multi-agent simulation framework evaluating LLM robustness to adversarial persuasion in simulated clinical emergency medicine scenarios.
Audit of 16 LLMs showing instability in ethical stances when moral dilemmas are reframed as negations versus prescriptions.
VERA-MH benchmark for validating AI chatbot safety in suicide risk detection with human evaluation studies.
Federated learning framework addressing device heterogeneity and non-IID data with adaptive differential privacy mechanisms.
Study showing sensitive attributes emerge in unsupervised embeddings even when withheld from training using self-organizing maps.
Theoretical analysis of aggregating multiple LLM responses in compound AI systems and whether this unlocks new capabilities.
Method for improving emotional reasoning in multimodal LLMs using reflective reinforcement learning for better emotion understanding.
Deep generative models for anomaly detection in multivariate time-series using normalizing flows with inductive biases in latent space.
Methods for measuring metacognitive capabilities of AI systems to assess reliability and manage uncertainty in decision-making workflows.
Research on smaller 4B parameter models for agentic execution tasks using subagent architectural patterns to handle specialized subtasks like debugging and terminal execution.
Study of how LLM-compressed financial summaries can distort investment decisions and information fidelity in agentic systems.
Analysis of how aligned LLMs internally represent safety through harmfulness and refusal directions for robust alignment.
Study of memory failures in LLM agents where conflicting state facts coexist, and methods to decouple them.
Stage-aware agentic framework for physical design optimization avoiding full re-runs after parameter changes.
CAD agent for industrial component design using knowledge distillation to handle ambiguous specifications and parametric modeling.
System for orchestrating mathematical reasoning agents using fact-graph memory for research-level problem solving.
Study of neural plasticity and generalization in deep reinforcement learning for adaptive video streaming.
Framework for Vision-Language-Action models reducing inference latency and improving robotic manipulation through trajectory ensemble voting.
Theoretical analysis of training solutions in two-layer neural networks with smooth activation functions.
Generative model using VAE with Bi-LSTM for time series data augmentation in forecasting and classification.
Study of visual reasoning about object properties and physical attributes in VQA tasks.
Survey on compositional visual reasoning in multimodal AI systems for decomposing scenes and multi-step inference.
Study of gradient-based jailbreak attacks on LLMs using adversarial suffixes without fixed target constraints.
Method to enhance text embedding models' semantic reasoning through multiple forward passes, tested on benchmarks.
Cross-task visual in-context learning via VLMs for handling mismatched demonstrations. Extends VICL to tasks differing from examples.
Benchmark and evaluation of foresight intelligence in vision-language models for anticipating future events. New VQA dataset for predictive capabilities.
Federated learning framework for multimodal feature extraction with contrastive learning under non-IID data. ML research with privacy focus.
Hierarchical Mixture-of-Experts framework for vision-language-action policies handling heterogeneous robot data. Enables generalist multimodal policies.
Compression technique for Mixture of Experts models using structured butterfly matrices. Reduces memory scaling for efficient edge deployment.
Token compression technique for omnimodal LLMs using audio-driven semantic chunking. Improves inference efficiency for multimodal models.
Framework for asynchronous multi-agent collaboration on long-horizon software engineering tasks. Addresses agent coordination and timely completion.
LLM-informed planning framework for object search in partially-known environments using prompt selection. Combines planning with LLM knowledge.