Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Beyond Leaderboards: Building the Evaluation Layer for AI
Learn how to build an evaluation layer for AI that treats results as reusable evidence, not just scores. Discover why standardizing data is crucial for debugging discrepancies and ensuring comparability.
We built connected open-source infrastructure for treating AI evaluation results as reusable evidence rather than isolated scores. Every Eval Ever standardizes results from harnesses, papers, leaderboards and local runs, while Evaluation Cards combines them with model and benchmark metadata and surfaces reproducibility, completeness, provenance and comparability. The live demo will trace a result from raw output through validation and normalization into an interactive card, then reveal why apparently identical scores can disagree.
Evaluation Cards reports AI model-benchmark results with four interpretive signals.
Every Eval Ever standardizes AI evaluation results via a shared schema and crowdsourced database.
- PythonPython: The high-level, general-purpose language built for readability, powering everything from web backends to advanced machine learning models.Python is the high-level, general-purpose language prioritizing clear, readable syntax (via significant indentation), ensuring rapid development for any team . Its ecosystem is massive: use it for robust web development with frameworks like Django and Flask, or leverage its power in data science with libraries such as Pandas and NumPy . The Python Package Index (PyPI) provides thousands of community-contributed modules, offering immediate solutions for tasks from network programming to GUI creation . The language is actively maintained by the Python Software Foundation (PSF), with the stable release currently at Python 3.14.0 (as of November 2025) .
- PydanticPydantic is Python's most-used data validation library: it enforces data schemas using standard type hints and boasts a Rust-core for exceptional speed.Pydantic is the premier data validation and parsing library for Python. It mandates data structure using pure, canonical Python type annotations, drastically reducing boilerplate code. With over 360M monthly downloads, Pydantic is battle-tested: all FAANG companies and major frameworks (FastAPI, SQLModel, LangChain) rely on it for robust data handling. Its core validation logic is written in Rust, ensuring high performance. Pydantic models also generate JSON Schema, facilitating seamless integration and documentation for API development.
- Hugging FaceHugging Face is the central, open-source platform and community for building AI applications, hosting over 300,000 models and datasets via the popular Transformers library.Hugging Face functions as the 'GitHub for machine learning,' providing a massive, collaborative Hub for AI assets (models, datasets, and demos). Its core technology is the open-source **Transformers** Python library, which simplifies the use of state-of-the-art models (e.g., BERT, GPT) for various tasks: natural language processing, computer vision, and audio. The platform hosts over 300,000 models and thousands of datasets, streamlining the entire ML workflow from research to deployment via **Spaces** (interactive demos). This ecosystem makes advanced AI accessible, efficient, and reproducible for developers and enterprises globally.
- DuckDBDuckDB is the high-performance, open-source, in-process analytical data management system (OLAP) that runs complex SQL queries directly on your data files.DuckDB is a fast, embedded analytical RDBMS: think SQLite, but optimized for OLAP workloads. It operates in-process—no separate server required—and boasts zero external dependencies, making deployment simple. The system uses a vectorized, column-oriented architecture for blazing-fast query execution on large datasets. It supports standard SQL, integrates seamlessly with languages like Python and R, and can query data directly from formats like Parquet, CSV, and JSON, eliminating ETL overhead. With over 25 million downloads per month and adoption by 20+ Fortune-100 companies, DuckDB is the go-to tool for local, high-speed data analysis.
- NextNext.js is the full-stack React framework: it delivers high-performance web applications via hybrid rendering and powerful, Rust-based tooling.This is the React Framework for production: Next.js enables you to build full-stack web applications with zero configuration and maximum efficiency. It supports a hybrid rendering approach (Server-Side Rendering, Static Site Generation, and Incremental Static Regeneration) for optimal speed and SEO performance. Key features include React Server Components, Server Actions for running server code directly, and the App Router for advanced routing and nested layouts. Developed by Vercel, it leverages Rust-based tools like Turbopack and the Speedy Web Compiler for the fastest possible builds and a superior developer experience.
Related talks
More from the community
From Abacus to AI: Small Potatoes, Big Results
Hong Kong
See how a non-coder built World of Warcraft addons, a bus app, and a phone ringer using AI…
EvoFit: Building a Cross-Cultural AI Fitness Coach That Bridges Eastern Wellness and Western Exercise Science
Hong Kong
Discover EvoFit, an AI fitness coach blending Eastern wellness and Western exercise science. See live demos of AI…
ai-flow.eu Evaluation Suite: Systematic Testing for LLM Apps
Cologne
Learn to systematically test LLM apps with ai-flow.eu. This demo covers building an evaluation suite, versioning datasets, running…
Culturally Aligned AI: Building Dlab-852-Mini for Hong Kong Cultural Nuances
Hong Kong
Showcasing Dlab-852-Mini, a Phi-3 fine-tune for Hong Kong culture using the CultureKit eval, detailing training, evaluation, and case…
Evaluating Multi-Agent Systems Beyond the Final Answer
Seattle
Learn how to evaluate multi-agent AI systems beyond just the final answer. This talk details a framework capturing…
Auto-create reliable LLM evals
New York City
Learn how to build reliable LLM-based evaluations using only about twenty human annotations, achieving scalable, human‑aligned assessments that…
Compose Email
Loading recent emails...