Automated generation of dynamical system computational models from natural language text using enhanced SysML diagrams and LLMs.
White-Basilisk proposes a hybrid model combining Mamba layers, linear self-attention, and Mixture of Experts for code vulnerability detection.
POLIS framework enables LLMs to accumulate knowledge through multi-agent interaction and inference, mimicking cumulative cultural evolution.
FLOSS framework enables federated learning with user opt-out and stragglers support, addressing data privacy in heterogeneous distributed systems.
ReasonRank method empowering passage ranking with reasoning ability by leveraging large reasoning models for improved listwise ranking in complex scenarios.
Establishes task-stratified knowledge scaling laws for post-training quantized LLMs, analyzing quantization impact on memorization, application, and reasoning capabilities.
SMARTER framework for explainable toxicity detection using LLMs with synthetic explanation generation and preference optimization for content moderation.
Studies transformer capability to learn transitive relation reasoning in graphs, essential for LLM factual correctness and causal inference tasks.
Multi-Level Optimal Transport (MOT) framework for aligning representational structures across model layers and brain regions with global alignment scoring.
Locate-Then-Examine two-stage VLM-based forensic framework for detecting AI-generated images using grounded region reasoning for artifact identification.
Reframes human label variation in NLP from noise to signal for improving model robustness, particularly relevant for post-training methods with human feedback.
Pilot study testing hypothesis that continual exposure to junk web text induces cognitive decline in LLMs using controlled Twitter/X corpus experiments.
CodeRL+ improves LLM code generation using reinforcement learning with execution semantics alignment, bridging gap between text patterns and functional correctness.
Analysis showing LLMs develop universal sinusoidal representations of numbers across different families, with representations largely interchangeable across models.
OpenHands Software Agent SDK provides a composable, extensible toolkit for building production-ready software engineering agents with flexible implementation and secure execution.
ItemRAG applies retrieval-augmented generation with item-based similarity to improve LLM-based recommendation systems, addressing cold-start problems.
AutoGraphAD uses variational graph autoencoders for unsupervised network intrusion detection without requiring labeled datasets.
Digital in-memory stochastic computing architecture using compressed bent-pyramid format to optimize AI model matrix multiplication operations.
Hybrid-AIRL method combining adversarial inverse reinforcement learning with supervised expert guidance, evaluated on poker game with imperfect information.
Research on sparse dictionary learning in mechanistic interpretability, analyzing how neural networks represent concepts as linear directions and encode multiple concepts.
Device-native autonomous agent system for privacy-preserving automated negotiations in insurance and B2B commerce without centralized servers.
CEDAR agentic system automating data science tasks via context engineering, handling complexity, data size, and computational constraints.
SciCoQA dataset of 635 paper-code discrepancies to evaluate LLM capability for cross-modal verification and research reproducibility auditing.
KOCO-BENCH benchmark evaluating how LLMs acquire and apply domain knowledge in specialized software development tasks.
Theoretical work relaxing realizability assumptions in language identification and generation tasks, establishing statistical rates without distribution constraints.
QuantaAlpha is an evolutionary LLM-driven agentic framework for financial alpha mining with multi-round search and experience reuse capabilities.
NeuroSymActive combines neural-symbolic reasoning with knowledge graphs for complex multi-hop question answering, integrating LLMs with structured knowledge.
CAST system for stable LLM-based text analysis of tabular data with algorithmic prompting to ensure output consistency for data analytics tasks.
Domain-specific LLM application for residential energy retrofit decision-making, guiding homeowners through building energy assessments.
Developer tool for flexible and high-performance fully sharded data parallel training, enabling block-wise quantization and structure-aware methods at scale.
Medical imaging application of vision-language models with pretraining for differential VQA tasks requiring fine-grained visual comparison.
Study analyzing semantic representation alignment between language models and vision encoders in vision-language models for taxonomic generalization.
Research on membership inference attacks against contrastive pretraining models like CLIP to audit PII memorization in multimodal backbones.
AutoML approach combining automated machine learning with deep unfolding for wireless beamforming optimization using learned proximal gradient descent layers.
Research evaluates robustness of climate foundation models under distribution shift, testing generalization beyond training data.
Explainable AI analysis revealing why AI-generated text detectors fail despite high benchmark accuracy, exploiting dataset-specific artifacts.
Visual Masked Autoencoder with Normalizing Flow for time series anomaly detection using foundation models for improved generalization.
CARLA-Air: Unified simulation infrastructure for air-ground embodied AI combining drone and vehicle dynamics in single environment.
Causal analysis showing LLMs fail reasoning when surface heuristics conflict with implicit constraints, studied through car wash problem.
Theoretical analysis of generalization bounds for overparameterized shallow neural networks related to distance from initialization.
Investigation of whether smaller language models with task-aware retrieval can match larger proprietary models for scientific knowledge discovery applications.
QuanBench+: Unified benchmark for evaluating LLMs on quantum code generation across Qiskit, PennyLane, and Cirq frameworks with 42 aligned tasks.
PromptEcho: Reward construction method using vision-language models for text-to-image model RL without human annotations or additional training.
Study on robustness of value alignment through finetuning with synthetic documents, releasing Animal Harm Benchmark for evaluating model compassion.
BenGER: Open-source web platform for end-to-end benchmarking of LLMs on German legal tasks with integrated annotation and evaluation workflows.
Comparison of LLM and human annotation in active learning for hostility detection on German political TikTok comments dataset.
Linear probing study analyzing how LLMs represent rhetorical questions in internal representations across social media discourse contexts.
Cognitive Reverse-Engineering framework for interpreting how LLMs internally process complex emotions and affective states through mechanistic analysis.
K-Token Merging method for compressing long sequences in LLM latent embedding space to reduce quadratic self-attention costs during inference.
LLM agent system for iterative data visualization refinement, automatically adjusting embedding algorithm configurations for high-dimensional exploratory analysis.