HN theanonymousone 21d ago

Grok 4.5 Benchmark Results

Benchmark results for Grok 4.5 proprietary LLM showing performance metrics, context window (500k tokens), and multimodal capabilities.

HN vinothkumarnaga 21d ago

Cybersecurity AI (CAI) Dataset

Cybersecurity AI (CAI) Dataset released for training and evaluating ML models on security tasks.

HN cooleel 22d ago

Accelerating Harbor with Tensorlake

Tensorlake runs Docker images for Harbor on microVM sandboxes, passing Terminal-Bench 2.1 tasks with cold starts in seconds.

Ax Sifat Afroj Moon, Dakotah Maguire, Adam Spannaus, Joe Tuccillo, Maksudul Alam, Sudip K. Seal, John Gounley, Heidi Hanson 22d ago

LLM-powered reasoning in agent-based modeling

Hybrid approach combining LLMs with agent-based modeling to enable real-time adaptive decision-making in large-scale individual interaction simulations.

Ax Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh 22d ago

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Analyzes token economics of enterprise agentic AI, arguing orchestration design (harness layer) is key lever against token inflation per task.

Ax Jerry Han, Rafael Moschopoulos, Ella Colby, Vishrut Goyal, Andrew Tu, Kia Ghods, Mark Braverman, Elad Hazan 22d ago

Measuring Intelligence Beyond Human Scale

Proposes relative-scale evaluation paradigm where AI models generate adversarial challenges to measure intelligence beyond human-saturating benchmarks.

Ax Elaine Ang, Chenxi Huang, Georgios Liargkovas, Jerry Liu, Jinhui Liu, Nikos Pagonas, Charlie Summers, Haonan Wang, Jiakai Xu, Tianle Zhou, Yusen Zhang, Zhou Yu, Zhuo Zhang, Tianyi Peng, Kostis Kaffes, Eugene Wu 22d ago

Agentic Data Environments

Agentic Data Environments framework for autonomous agents operating across files, APIs, applications, and system state with failure bounds.

Ax Adam Jenkins, Agnieszka Kitkowska, Caterina Maidhof, Diego Paracuellos, Francesco Sovrano, Gonzalo Gabriel Mendez, Guillermo Suarez-Tangil, Hana Kopecka, Isabel Wagner, Isabel Barbera, Javier Carnerero-Cano, Jide Edu, Jose Luis Martin-Navarro, Jose Such, Josep Domingo-Ferrer, Juan Carlos Carrillo, Kopo Marvin Ramokapane, Mark Cote, Pablo Vellosillo, Ramon Ruiz-Dolz, Rongjun Ma, Ruba Abu-Salma, Sameer Patil, William Seymour, Xiao Zhan 22d ago

Security and Privacy in Agentic AI: Grand Challenges and Future Directions

Horizon-scanning study identifying security and privacy challenges in agentic AI systems.