AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
AMALIA: fully open-source LLM trained on high-quality European Portuguese data with native evaluation benchmark.
AMALIA: fully open-source LLM trained on high-quality European Portuguese data with native evaluation benchmark.
JAL-Turn: joint acoustic-linguistic model for turn-taking detection in real-time voice AI agent systems.
ALBA benchmark for evaluating European Portuguese language understanding in generative LLMs, addressing underrepresented language evaluation.
Analysis of how open vs closed language models impact scientific inference reliability and reproducibility.
Analysis of efficiency metrics for vision backbone networks, demonstrating MACs limitations on edge devices versus actual execution time.
Study on knowledge distillation of Transformers into efficient hybrid models with focus on generation quality beyond perplexity metrics.
StackRepoQA: first repository-level benchmark evaluating LLMs on multi-file code question answering and program comprehension.
Framework profiling energy, latency, and quality trade-offs of running LLMs on edge devices with memory and thermal constraints.
Study on improving vision-language models' spatial reasoning by injecting geometry tokens from 3D foundation models.
Vision2Web: hierarchical benchmark for evaluating AI agents on website development tasks from UI-to-code generation through full-stack development.
PerceptionComp: benchmark for complex long-horizon perception-centric video reasoning requiring compositional temporal evidence and multi-subtask integration.
Ruka-v2: open-source tendon-driven dexterous humanoid hand with 11 DOF, data-driven finger control, buildable under $1,300.
Scale-adaptive exploration-exploitation balancing in classical planning using Multi-Armed Bandit theory for improved MCTS-based planners.
Extreme value Monte Carlo Tree Search improves unbounded cost-to-go estimation in classical planning, outperforming UCB1-based approaches.
ReMe: LLM-mediated conversational framework for scalable personalized cognitive training with controllable task structure and natural interaction.
Deontic temporal logic formalization for specifying and verifying ethical behavior in AI systems through formal methods.
ProbGuard: probabilistic runtime monitoring framework for LLM agent safety enabling proactive risk detection during execution in robotics/automation.
Online alignment (GRPO) outperforms offline alignment (DPO) through better approximation of human-perceived model output distribution based on prospect theory.
Causal perspective on reasoning tasks showing selection and reflection mechanisms improve LLM reasoning reliability through self-refinement approaches.
Multi-agent predictive coding framework for constructing shared spatial memory through mutual uncertainty minimization with information bottleneck objective.
HeaRT: hierarchical circuit reasoning agentic framework for analog/mixed-signal design optimization with adaptive mechanisms and improved transferability.
Analysis of decision-making failures in foundation models for navigation across reasoning tasks with complete/incomplete spatial and safety-relevant information.
AtomMem: learnable dynamic memory framework for agents using atomic operations, replacing static hand-crafted workflows for long-horizon problem solving.
AIDABench: comprehensive benchmark for evaluating AI-driven document understanding and processing tools in end-to-end real-world scenarios.
Draft-and-Prune method improves auto-formalization reliability by reducing semantic failures when translating natural language to solver-executable logical programs.
NLP study detecting political propaganda on Moltbook platform using LLM-based classifiers on 673k posts; analyzes prevalence and concentration patterns.
Governance-aware vector subscriptions for multi-agent systems enabling secure real-time knowledge monitoring with data handling policy enforcement.
Environment Maps: persistent agent-agnostic representation to reduce cascading errors in long-horizon LLM-based automation tasks across dynamic interfaces.
Formal framework analyzing whether AI assistants improve or introduce blind spots in safety engineering workflows for physical AI systems.
Research paper proposing deep symbolic regression approach using policy gradients for discovering interpretable mathematical expressions from data.
Open-source CGRA4ML framework for implementing neural networks on configurable hardware for near-sensor scientific edge computing.
Research paper proposing INSIGHT framework using vision-language models for hazard detection and edge case evaluation in autonomous driving.
Research paper proposing Biogeochemistry-Informed Neural Networks integrating domain knowledge with ML for soil carbon modeling.
DARai: multimodal hierarchically-annotated dataset of 200+ hours of daily human activities across 10 environments with 20 sensor types.
Research paper proposing FastCache framework for accelerating Diffusion Transformer inference through learnable linear approximation and token caching.
Multi-agent copilot using multimodal LLMs for evidence-based diagnostic reasoning in digital pathology.
StreamDiT enables real-time streaming text-to-video generation using transformer-based diffusion models.
LLM-based framework using chain-of-thought and reinforcement learning for interpretable cyclic peptide design.
3D Gaussian Splatting method for open-vocabulary scene understanding by decoupling geometry and semantics.
arXiv research on ATAR method improving LLM reasoning by aligning attention with reasoning structure. Prevents critical steps from being buried in extended reasoning chains.
arXiv research on unified speech and gesture synthesis via interleaved token prediction. Jointly generates synchronized co-speech gestures and speech from text with discrete autoregressive model.
arXiv research auditing racial and gender biases in LLM-generated occupational personas across 41 professions. Analyzes 1.5M personas from GPT-4, Gemini, DeepSeek, Mistral against BLS data.
arXiv research on training-free compositional image synthesis combining object-centric approaches with self-refinement. Uses LLMs to improve layout faithfulness in text-to-image generation.
arXiv research on GUI grounding for computer-use agents. Proposes intrinsic multimodal attention alignment for mapping natural language to screen regions instead of coordinate generation.
arXiv research on Sequence-level TopK routing for Mixture-of-Experts LLMs. Minimal modification enabling adaptive expert routing based on token complexity without retraining.
arXiv research on training-free binary verification workflow for zero-shot vision using VLMs. Converts open-ended queries to multiple-choice with deterministic resolution.
arXiv research on 4D generation from natural language and images using embodied world models. Addresses data scarcity and long-horizon video generation challenges.
arXiv research proposing Balanced Fine-Tuning method for aligning LLMs with biomedical knowledge. Combines SFT and RL using confidence-weighted token optimization for scientific understanding.
arXiv research on streaming video understanding with gaze signal interpretation for AR applications. Evaluates multimodal LLMs on temporal reasoning with human attention signals.
arXiv research on multimodal memory architecture for long-form video understanding. Addresses context capacity and visual detail retention in hours-long videos using dynamic memory mechanisms.