1
1uc4sm4theus
Guest
A knowledge-graph investigation tool built on Nosana's decentralized compute network, born inside a small Brazilian recommendation-tech company.
A scandal too big to read
In November 2025, Brazil's Central Bank ordered the extrajudicial liquidation of Banco Master, a mid-sized lender that had grown explosively by selling roughly R$50 billion in bank deposit certificates (CDBs) — instruments backed by the government's deposit guarantee fund, but themselves propped up by illiquid, hard-to-value assets like court-ordered payment rights (precatórios) and stakes in struggling companies. Days later its controller, banker Daniel Vorcaro, was arrested on money-laundering charges. Brazilian press quickly began calling it the largest banking fraud in the country's history.
)
What made the case explode beyond a financial story, though, wasn't the accounting. It was the release of leaked message threads — later partially unsealed by the Federal Police and reported on by outlets like Aos Fatos and ICL Notícias, O Globo — appearing to show Vorcaro in direct contact with Central Bank staff, judiciary members, and political figures across the ideological spectrum, allegedly trying to soften or delay regulatory action against his bank. Within weeks, dozens of names, institutions, and message threads were circulating across news sites, PDFs, and social media, each outlet publishing a different fragment of the same underlying network.
/https://i.s3.glbimg.com/v1/AUTH_da025474c0c44edd99332dddb09cabe8/internal_photos/bs/2026/2/s/0xZQUVRjGGl3T2c4MTrA/pol-27-09-mensagens-vorcaro.png)
O GLOBO'Concierge' do poder: mensagens obtidas pela PF mostram como Vorcaro aprimorou método usado para influenciar autoridadesAvaliação de profissionais do Direito e da ciência política é que o banqueiro se especializou em promover “fantasias”, “ostentação” e “tentações exuberantes”
That's a hard shape of problem for the public to hold in their heads. A financial fraud with an evolving cast of forty-plus people, examined through hundreds of individual messages scattered across dozens of articles, is exactly the kind of information environment where confusion, rumor, and cherry-picked screenshots thrive instead of understanding. So at RecomendeMe — a Brazilian company that already applies graph methodology to culture and research through its RecomendeMe Intelligence arm — we asked a narrower question: could we turn the public record of this case into something a citizen, or a journalist without a data team, could actually query?
That question became MasterWhats.
What MasterWhats actually is
MasterWhats is a web application designed to make the Vorcaro case accessible to the public in a structured, readable way. It organizes people, institutions, messages, and source documents into an interface where anyone can explore the case and understand the evidence without having to navigate hundreds of documents.
Behind the application is a GraphRAG pipeline: a structured knowledge graph connected to a retrieval-augmented generation system that answers plain-language questions and points back to the specific dated message and source document behind each answer.
The graph currently contains 49 nodes, 29 relationship edges, and 241 individually indexed evidence chunks. Each piece of evidence remains linked to its original message, date, and source document, preserving the context needed to distinguish documented connections from assumptions or coincidences.
MasterWhats is therefore not a leak database and not an accusation engine. It is an interface for exploring and understanding a complex public case through structured, traceable evidence.
The pipeline, step by step
The notebook that runs MasterWhats does four things, in order:
1. Build a citable evidence corpus. For every edge in the graph, it unpacks the evidence array attached to it and produces one text record per message: something like "[date] Person A → Person B (action): quoted message text (source: official document reference)". Edges with no granular evidence still get indexed as a bare fact, but are clearly distinguishable from cited claims.
2. Embed everything for semantic search. Each evidence record is embedded with
intfloat/multilingual-e5-large, a multilingual sentence-transformer model that performs well on Brazilian Portuguese — important, since the underlying messages are informal, code-switched, and full of local slang — and indexed into a FAISS flat index for fast similarity search.3. Ground an LLM strictly in retrieved evidence. A quantized Qwen2.5-3B-Instruct model — small enough to run affordably, large enough to reason over Portuguese text — answers questions by first retrieving the top-k most relevant evidence chunks, then generating an answer built only from what was retrieved. The prompt explicitly instructs the model to say when the evidence is insufficient rather than fill gaps with speculation. That instruction is the whole point: the system is designed to refuse an answer it can't source, not to produce a confident narrative.
4. Attempt to resolve unknowns, and generate per-person dossiers. For nodes still marked
UNKNOWN, the pipeline gathers all the surrounding connections and asks the model for a hypothesis with an explicit confidence level — low, medium, or high — rather than a flat assertion. And for any named person, it can generate a dossier built exclusively from evidence directly connected to that individual's node, with an instruction to explicitly flag thin or ambiguous evidence instead of stretching it into a narrative.The output of every query is traceable: ask "who was in contact with whom about the Central Bank's monitoring flags," and the answer comes back cited to a dated message and a document reference, not a vibe.
Why it runs on Nosana
None of this is exotic machine learning — a 3B-parameter instruction model and a multilingual embedding model are well within reach of a single consumer GPU. But "well within reach of a GPU" and "well within reach of a small, self-funded team in Natal, Brazil" are two different things. RecomendeMe has grown entirely without outside investment.
That's where Nosana came in. Nosana is a decentralized GPU marketplace built on Solana: instead of renting fixed capacity from a hyperscaler, it connects idle GPUs — from data centers, gaming rigs, and former crypto-mining hardware around the world — to people who need on-demand inference compute, coordinated and paid for on-chain. For a project like MasterWhats, that model is a good structural fit in a few concrete ways:
- No long-term commitment for a bursty workload. Embedding 241 evidence chunks and running batches of grounded Q&A doesn't need a GPU sitting reserved 24/7 — it needs one available when a Jupyter session spins up, and Nosana's node marketplace is built exactly for that kind of on-demand, pay-for-what-you-use pattern.
- Cost proportional to an independent, bootstrapped project, not to enterprise cloud pricing.
- A civic-tech use case that fits the network's own thesis — turning otherwise-idle compute around the world into infrastructure for something other than another commercial inference API, in this case a public-interest investigation tool.
Running on distributed, permissionless GPU infrastructure also has a quieter symbolic fit: a project meant to make an opaque, power-adjacent financial scandal more transparent runs on compute that nobody in Brasília controls the switch to.
What it actually changes for the public
The Banco Master case has been extraordinarily well covered by Brazilian fact-checkers and investigative outlets in particular, and has done the hard work of verifying and contextualizing leaked messages as they surface. MasterWhats isn't trying to replace that reporting or compete with it. It's trying to solve a different, more structural problem: once a dozen outlets have each published their own fragment of the network, there's no single place for someone to ask "how are these two people actually connected?" and get an answer that (a) is grounded in a specific citable source, (b) tells you plainly when the evidence doesn't support a conclusion, and (c) doesn't require you to have read forty articles first.
That's a meaningfully different kind of access than a search engine or a single explainer article provides. It's closer to what investigative newsrooms build internally for large document leaks — a queryable, cited index — made available as a public tool instead of a private reporting aid.
It also comes with real limits we're upfront about. The graph reflects what has been publicly reported and unsealed, not the full case file. The
UNKNOWN-node resolution and per-person dossiers are explicitly framed as hypotheses with confidence levels, not verdicts — nobody is "guilty" because a language model connected two names. And a 3B-parameter model, however carefully grounded, can still misread nuance in Portuguese slang or informal phrasing, which is exactly why every answer ships with its source attached: so a human can check it.Where this goes next
The notebook's own roadmap points at the obvious next steps: swapping the local FAISS index for Neo4j's native vector index — RecomendeMe Intelligence's existing OSINT stack already runs on Neo4j — so semantic search and graph traversal (Cypher queries) can be combined in one system; adding a confidence-auditing pass where the model re-checks each edge's stated confidence against its cited evidence; and exposing
ask_graph() behind a simple API to power a public "ask the graph" feature on masterwhats.recomendeme.com.br.The broader bet behind MasterWhats is that as financial and political scandals keep generating this same shape of problem — sprawling, leak-driven, too large for any one article to hold — small teams with a graph, a modest open model, and access to affordable decentralized compute can build public infrastructure for understanding them, without needing a newsroom's budget or a cloud contract to do it.