PEARL: Personalized Streaming Video Understanding Model
PEARL: streaming video understanding model for personalized real-time interactive responses from continuous visual input and instant feedback.
PEARL: streaming video understanding model for personalized real-time interactive responses from continuous visual input and instant feedback.
Study showing coding agents effectively process long context by externalizing processing into explicit executable interactions rather than latent attention.
ALICE evaluation framework for assessing in-context learning ability of six large audio-language models under audio conditioning with reduced guidance.
Melaguard: lightweight Transformer edge AI framework (1.2M parameters) for detecting neurovascular instability from wearable physiological signals.
Solver-aided verification framework ensuring tool-augmented LLM agents comply with domain-specific operational policies for sensitive applications.
Study showing LLM-detection policies for peer review are unenforceable; evaluates five detectors on human-AI collaboration scenarios.
Feature attribution method for ECG signal analysis to explain machine learning model decisions in biomedical applications.
Diffutron: masked diffusion language model for Turkish using LoRA-based continual pretraining, addressing morphologically rich language modeling.
Application of emotion AI and intercultural pragmatics to measure affective engagement in language learning contexts.
Co-design approach for fully homomorphic encryption and AI inference to enable privacy-preserving cloud model execution without weight updates.
Framework for evaluating legibility of reasoning traces in LLMs that output chain-of-thought deliberations alongside final answers.
ReBOL retrieval system using Bayesian optimization and LLM-based reranking to overcome vector similarity limitations and enable query reformulation.
Evaluation of GPT-4, Gemini Pro, Llama-3, Mistral-7B on health crisis queries in low-resource Bangladesh context using multi-metric assessment.
Distributed reinforcement learning approach addressing negative learning from high-surprisal data in stale or mismatched actor scenarios.
Policy gradient optimization method using 'delight' metric to selectively compute backward passes, reducing computation while maintaining learning value.
Analysis of inverse correlation between model confidence and accuracy across Llama, Qwen, Mistral, and OLMo; formal proof that miscalibration is observational not capability gap.
Business model analysis of generative AI platforms exploring multi-layer market architectures and revenue-sharing infrastructure for API/model providers.
Industrial-scale RAG framework for requirements engineering in automotive manufacturing, with empirical evaluation on production data and performance metrics.
PCFJudge: Inference-time method using permutation consensus to improve robustness of LLM-based factuality judges against candidate order sensitivity.
MKA: Memory-Keyed Attention mechanism for efficient long-context LLM inference by reducing KV cache size while maintaining representation quality.
AEGIS: Graph-guided LLM agent for vulnerability detection using dialectics and meta-auditing to ground reasoning in code-specific evidence.
Psychophysical analysis of how transformer language models represent magnitude using representational similarity, discrimination, and causal intervention.
Multihead continual learning framework for fine-grained fashion image retrieval using contrastive learning and exponential moving average distillation.
Study on how AI scaling laws and specialized accelerators reshape classical Amdahl's Law for modern heterogeneous computer architecture.
REVERE: Reflective evolving research engineer agent that optimizes prompts and workflows using broader task patterns for scientific research coding.
PAVE: Inference-time validation layer for retrieval-augmented LLMs that decomposes context into atomic facts to verify answer support.
SWE-Next: Execution-grounded framework for scalable software engineering task collection to train AI agents on real-world code repositories.
Network-of-Thought models LLM reasoning as directed graph instead of linear or tree structures for complex multi-source reasoning.
DiT-BlockSkip reduces memory during diffusion transformer fine-tuning through dynamic patch sampling and block skipping.
Compass optimizes compound AI workflows orchestrating multiple specialized models for accuracy, latency, and cost under dynamic loads.
MERIT addresses domain shifts in RAW images across camera sensors via efficient multi-domain translation model.
Dodgersort uses CLIP-based ranking with uncertainty decomposition and information-theoretic sampling to reduce pairwise comparison labeling cost.
HiCI proposes hierarchical attention module for long-context language modeling using segment-level representations and global integration.
Evaluates ChatGPT's comprehension of modern Chinese poetry through framework developed with professional poets.
Develops SozKZ, family of 50M-600M parameter Llama models trained from scratch on Kazakh with dedicated tokenizer.
Addresses neural network saturation during transfer learning by restoring plasticity of pretrained weights for downstream tasks.
Introduces semantic sections as feature ontology for interpreting neural networks in obstructed representation spaces.
RubricRAG generates interpretable rubrics for LLM evaluation using domain knowledge retrieval to replace opaque scalar scoring.
Proposes manifold-constrained hyper-connections to stabilize training of models with multi-stream residual architectures.
Applies natural gradient descent to online continual learning for image classification to prevent catastrophic forgetting.
Proposes SART, a gradient-aware training method to detect and mitigate shortcut reasoning in LLMs through gradient surgery and scoring.
Enhances LIME interpretability framework using neural decision trees for better explanation of complex ML models on tabular data.
Benchmarks deep learning models across GPUs to study computational efficiency and access disparities in training large-scale models.
CGDFS method for stable feature selection using causal principles to identify features robust to distribution shifts.
AC4A access control framework enabling fine-grained permission management for LLM agents interacting with external APIs and tools.
VARS framework for persistent user preference modeling in conversational LLM agents using retrieval-augmented interaction.
Deterministic pre-action authorization framework for AI agent tool calls enforcing permission-based policy before execution.
Contrastive learning approach for discovering functional gene associations from protein interaction networks.
Demonstrates finetuning bypasses safety alignment in LLMs, activating verbatim recall of copyrighted training data.
Multi-agent framework aggregating zero-shot LLM judgments for corporate disclosure classification using trained lightweight aggregator.