Daniel Alami
Researching how to verify AI-generated work.
I study the evidence behind AI systems: whether generated tests distinguish faulty code from correct code, whether benchmark results support a model choice, and whether an agent's sources justify its conclusions. I build public tools to make those checks easier to inspect and repeat.
At Amazon Web Services, I am a Senior Product Manager - Technical working on neurosymbolic AI systems and evaluations for production-grade product and pricing authoring. The research and software below are part of my independent public work.
I hold an MSc in Information Science from Utrecht University and an MBA with Distinction from Harvard Business School. My earlier roles included IBM, Gartner, Microsoft, and Santander.
Research
My earlier studies measured the health of software ecosystems, examined how people learn security requirements, and made institutional records usable as linked data. Across those projects, the common task was to turn broad claims about a system into something others could examine. That is also the thread in my current AI research.
I now work on three related problems: evaluating AI-generated tests without rejecting valid solutions; checking when model comparisons remain stable enough to guide a decision; and deciding who may approve changes to the evaluators and policies that guide AI agents. The goal is practical methods that other researchers and teams can reproduce, challenge, and use.
Peer-reviewed papers
Three papers have been published and cited. Two further papers have been accepted for poster presentation at NeurIPS 2026 workshops: Trust-AI-Eval and Meta-Agents. Both manuscripts are awaiting public release. My Google Scholar profile has the citation record for the published work.
- A Gamified Tutorial for Learning About Security Requirements Engineering. Daniel Alami and Fabiano Dalpiaz. IEEE International Requirements Engineering Conference, 2017.
- Relating Health to Platform Success: Exploring Three E-commerce Ecosystems. Daniel Alami, Maria Rodríguez, and Slinger Jansen. European Conference on Software Architecture Workshops, 2015.
- Experiences from the Design and Development of an Institutional Linked Open Data Portal. Daniel Alami, Isaac Lera, Carlos Guerrero, and Carlos Juiz. TEM Journal 6(4), 2017.
- What Does Fifty Days Stabilize? A Commit-Level Audit of ForecastBench. Daniel Alami. Accepted September 21, 2026, for poster presentation at the NeurIPS 2026 Trust-AI-Eval workshop; manuscript pending public release.
- The Cognitive Firm: Authority, Independence, and Control in Organizations of Humans and AI Agents. Daniel Alami. Accepted September 30, 2026, for poster presentation at the NeurIPS 2026 Meta-Agents workshop. A position paper on who may approve evaluator repairs, policy changes, and delegation changes in AI-agent workflows; manuscript pending public release.