lucas buchalla sestivalencia, spain +34 666 277 651rev. 2026 — applied ai / mobile systemsen/es

Product
systems
engineer

flutter in production. 20k+ users.
agent platform in production. two tenants.

I build and run a multi-tenant AI analytics platform that two companies query against their own data, and I own the mobile architecture behind one of them. The same architecture rules apply to both.

langgraph supervisor, five specialist agents,
ground-truth evals, traces on every step.

2 tenants
one agent platform, isolated
5 suites
ground-truth evals, human-judged
20k+
mobile users, 97% crash-free
8+ years
intern → senior

What I do,
plainly

Lucas Buchalla Sesti

I build and operate an AI analytics platform: a LangGraph supervisor that routes to five specialist agents over BigQuery, a code graph of the repositories, a per-tenant RAG index and mobile vitals. Two companies run on it, each against its own data, fully isolated.

I also own the mobile architecture at Lejour: Flutter, TDD, design systems, 20,000+ users at 97% crash-free. The two jobs share a stack: the crash reports and BigQuery tables the agents read are ones I already maintain.

Evaluation is where I spend most of my time. Every change runs against five ground-truth benchmark suites with human judging and a trace on every step, and I keep the negative results as carefully as the wins. Most of the computer vision work below is negative results.

measure it · ground every claim · keep the negative results

Applied AI

two systems in production, one research track, all measured

Multi-tenant
AI Analytics

in production

Business and engineering questions asked in plain language, answered across four sources that never talked to each other before. Every answer comes back with the query, the commit and the trace behind it.

bigquery
product data
codegraph
repositories
chromadb
rag corpus
vitals
asc + play
supervisor agent — langgraph · adaptive reflect loop · gemini on vertex ai
↳ "why did crash-free drop on android last week?"
  vitals for the drop, then the commits that shipped that week.
problem

Answers lived in four silos. "Is this regression ours?" cost a day of three people’s time, every time it was asked.

architecture

A LangGraph supervisor with an LLM reflect loop that picks the next specialist from the evidence collected so far: metric, dimensional, code, code-review, mobile. CodeGraphContext over the repos, ChromaDB per tenant, BigQuery for product data, App Store Connect and Play Developer Reporting for vitals. Semantic-model fast paths answer common questions without an LLM in the SQL path. Provider-agnostic LLM factory, Postgres investigation store, tenant isolation from the first commit.

outcome

Two companies query one engine against their own data, fully isolated. When the evidence conflicts it says so and shows both hypotheses instead of picking one. The standing rule is that it never returns a causal claim it cannot ground in a source.

Football
computer vision

independent research
research track

Player re-identification and pose recognition on full-match footage. The deterministic tracker fragments 17 real people into 415 IDs across 55 minutes; holding identity through occlusion, camera cuts and near-identical kits is the open problem.

Fourteen measured phases so far. A fine-tuned ReID network moved pair AUC from 0.913 to 0.952 and still did not solve it, because it trains on contaminated labels: a third to a half of raw tracks swap person mid-track. The virtual camera did land. Non-causal smoothing and mode-based aiming cut direction reversals from 25.1/min to 3.7, and aim error from 653px to 457px against 18 hand-labelled positions. Every dead end is written up with its numbers so it never gets run twice.

reid   pose estimation   yolo   pytorch   opencv

read the full research log

How I know it works

evaluation harness

Five ground-truth benchmark suites, covering metric, code, RAG, threading and end-to-end, run against a seeded tenant with planted commits and indexed docs. Runs are labelled, replayed to measure variance, judged by hand and compared suite to suite, with a trace on every step. No mocks in the normal suite: it runs against real BigQuery and a real model. It is the least glamorous part of the platform and the reason I trust what it says.

Agent delivery
workflow

internal tooling
in production

A state machine that governs how a coding agent ships changes in a 3,500-commit Flutter repo. It routes the request, walks the task through spec, implementation and verification, and refuses to mark anything verified on the agent’s word.

The router scores the raw prompt against 23 weighted regex signals in Portuguese and English to pick one of five modes. The release gate then runs the test suite and the analyzer itself and reads the real counts: red tests, analysis above baseline, code touched with no test, code touched with no doc, or an empty diff each block the task, and the evidence lands in a JSON file per task. A context gate keeps 13 invariants, each scoped to a path glob, so a change pulls in only the rules covering the files it touched. The gate once passed green without analysing anything: it called the analyzer with a flag that version did not have, which exited printing its usage, and the usage contained the one word the parser searched for. It now requires the analyzer’s own summary line as proof it ran.

dart cli   state machine   release gates   context engineering

Track record

eight years, three products, one open-source package

Lejour

senior mobile developer02/2019 — presentremote (br hq), from spain

Wedding platform: couples build a site and a gift registry, guests RSVP and send photos, and professionals manage everything from a separate app. Three Flutter apps on a shared internal library, where I own the mobile architecture, the design system and the testing culture, and I built and run the analytics platform above.

−25%
codebase
−60%
build time
−40%
dev time, design system
3+
devs mentored, one intern→mid

Mude Health

mobile developer02/2025 — 07/2025part-time, pt hq

Multi-brand white-label wellness apps in Flutter: one shared codebase, per-brand theming and releases, and a component library reused across deployments.

Onboardly

open source — pub.dev

Flutter package for onboarding flows: spotlight overlays and interactive tooltips, zero external dependencies, state-management agnostic. 160 pub points, ~197 weekly downloads, five platforms.

ai & data

LangGraph · supervisor + reflect loopsVertex AI · Gemini · provider-agnosticRAG · ChromaDB · semantic modelsBigQuery · PostgreSQL · Python

mobile

Flutter · DartSwift · KotliniOS · AndroidDesign systems

backend & infra

NestJS · Node · TypeScriptDocker · AnsibleGCP · AWS · FirebaseSupabase · Postgres

practice

Evals · benchmarks · tracingTDD · unit testingArchitecture · modularityMentoring · code review

languagesportuguese native · english c2spanish b2 · chinese a1 · arabic a1
certificationschinese — fluency academy, 2023arabic basic 2 — centro da língua árabe, 2023
built mobile-firstthis page was laid out at 375pxbefore it was laid out at 1440.