A generalist agent built this
Architecture
Two companion agents
The explorer now keeps two analysis paths side by side. Google
Conversational Analytics uses the published molto_logs_agent to resolve
schema, run BigQuery SQL, and return tables and charts. Codex +
OKF runs codex exec --model gpt-5.6-luna and consumes the OKF wiki from
the live Google Dataplex Knowledge Catalog EntryGroup molto_logs_okf.
Codex + OKF can request one live query per turn, but it never receives Google credentials or direct
network access. The Flask app parses generated GoogleSQL, accepts only one read-only
SELECT whose fully qualified sources are the four privacy-filtered public views,
dry-runs it with a bytes cap, and executes it using the dedicated
ca-luna-query service account. DDL, DML, wildcards, raw tables, other datasets,
and multiple statements are rejected before BigQuery sees them. Results are capped and returned
to Codex + OKF as explicitly untrusted data for a final answer. Separately, Flask pulls and translates
the Dataplex catalog, caches it for five minutes, and includes it in the prompt. Codex + OKF's sandbox
network remains disabled and API-key variables are stripped from child commands.
Open Knowledge Format
Open Knowledge Format (OKF) is a portable wiki convention: one concept per Markdown file, YAML frontmatter for compact structured signals, ordinary links for relationships, and paths as concept identities. An agent needs no special SDK to consume itβit can read the directory like any other source tree.
The OKF wiki published to molto_logs_okf documents the dataset, raw tables, public views, field
meanings, join cardinalities, query practices, privacy boundaries, and the known stale-schema
history of the CA agent instruction. Codex + OKF consumes the catalog copy; the repo's okf/
directory remains the authoring source used by the separate publish workflow. The wiki uses the
v0.1 signal fields resource,
type, generated, and sources alongside title,
description, and tags. The open specification and examples live in Google's
knowledge-catalog repository.
Regenerate and publish the wiki
Schema prose should follow live BigQuery evidence, never a model's remembered schema. The
refresh script queries INFORMATION_SCHEMA.COLUMNS,
COLUMN_FIELD_PATHS, and TABLES, plus table metadata, without reading
row values. Review its JSON before updating the table concept files.
Google's toolbox/mdcode is vendored at a pinned upstream commit. Its OKF demo
adapter stages the clean Markdown signal layer into a custom okf Dataplex aspect,
while the Documents Layout stores titles, descriptions, tags, and bodies. Publishing creates
EntryGroup molto_logs_okf in global.
The final push requires roles/dataplex.catalogEditor on project
moltonhim-agent for
ca-explorer-app@moltonhim-agent.iam.gserviceaccount.com. The app never grants IAM
itself. Full commands and round-trip pull instructions are in the vendored demo README.
What's interesting
Molto is an AI agent that runs 24/7 on a GCP VM, handling conversations, tasks, and self-maintenance. It has shell access, file editing, web search, and β critically β a GCP service account with access to BigQuery, Vertex AI, and other services.
When asked to build an analytics dashboard for its own operational logs, the agent:
agent_turns_full table in BigQuery containing
turn-level telemetry: tokens, latency, tool calls, errors, sender/channel metadata.
agent_turns_public β a BQ view that strips all message content
(user questions, agent responses, tool I/O) while preserving operational metrics.
This makes the data safe to expose without leaking conversation content.
The data agent resource
This is the live configuration that powers the analytics chat. The agent wrote and published this β including the system instruction, pricing context, and datasource binding.
Key details
moltonhim-agentmolto_logs_agentagent_turns_public + 3 current-pipeline public viewsWhy this matters
This wasn't built by a purpose-built data tool. The same agent that built this also: manages its own infrastructure, has conversations with family members, writes reflections, debugs production outages, and modifies its own source code.
The data agent β the Gemini Analytics resource that handles the NLβSQLβchart pipeline β was created and configured by a generalist agent that happened to have GCP access. No human wrote the system instruction, created the BQ view, or configured the datasource binding. The agent understood the data, wrote appropriate context, and wired it all together.
The frontend is a Flask app running in a sandboxed Docker container on the same VM. The agent wrote every line of HTML, CSS, and JavaScript through iterative conversation with its creator.