BioCreative

The Operating System

How BioCreative builds — for itself and for clients. The method is contextual engineering: assembling the right context, skill, tools, and harness for one specific task, at the right level of autonomy, in a system that writes its results back into the same modular context it read from. This page is the launcher; each section drills into the real artifact.

Section 1

The Agentic Operating System — Software Around the Processor

This section is the quick read of the whole thing; every piece below goes deeper. The one idea underneath all of it: the LLM is a processor, and from a single chat all the way up to a complex fleet of harnesses, it is always the same processor. Our work is to standardize the pieces that come together around it — context, skills, tools, and the harness — so we can assemble the right thing for the right agent, on the right surface, at the right time, with persistent data, skills, rules, and context wrapped around everything. The buckets are ours and the processor is swappable, so as the models get better our whole operation gets better without a rewrite.

The unit — how one agent is assembled
Context
What it must know. Authored in git, projected into SQL per tenant.
+
Skills
How to do it. 220 of them, composable atom → bundle → workflow.
+
Tools
What it acts on. 83 registered surfaces with credentials and bills.
→
Agent
The processor wrapped in a harness, at the right autonomy for the job.
→
Task, done
A real outcome, not a general answer.
↵  And the context is continually refined. The loop is the point: the shared context every agent reads from gets better over time. The primary path is backend refinement through git — humans editing the knowledge graph, projected into SQL (the Spine, §3). The second, emerging path is memory — learning events from the running fleet that eventually roll back up into the graph (Retrieval & Graph, §7). A system that consumes context but never refines it degrades; one that does compounds.
The scale — one agent becomes the system
One agent
The unit above: pieces assembled around the processor to finish one kind of task.
→
A coordinated fleet
80+ agents sharing one skills library, one brand bible, one contextual spine — handing work between each other.
→
The operating system
Built and coordinated right, the pieces compound into an operating-system-grade capability that optimizes across GTM, delivery, and operations.
The pieces — everything below, in one view
§2
LLM as Processor
The core. Six altitudes from a single chat up to an agentic loop — the same processor, more scaffolding.
Jump ↓
§3
Modular Contextual Spine
The shared context every agent reads from — a Karpathy/Wikipedia-style knowledge graph in git, projected into a structured SQL layer for downstream ops.
Jump ↓
§4
Skills
The composition ladder — 220 reusable procedures, atom → bundle → workflow, brand- and tenant-agnostic.
Jump ↓
§5
Tools
What agents act on — 83 registered surfaces, subscriptions, and BC services; the custom interconnection is the edge.
Jump ↓
§6
Agents, Harnesses & the Fleet
How one agent is built (the eight slots) and how 80+ coordinate across 27 harness patterns.
Jump ↓
§7
Retrieval & Graph
How the right context is pulled at query time — the honest state of what we run and what we don't yet.
Jump ↓
§8
UX & Operator Surfaces
The eyes — the dashboards and consoles where humans read the system and steer it.
Jump ↓

None of these is the product on its own. Coordinated — the spine feeding the skills, the skills composing into agents, the agents running in harnesses, and the shared context continually refined beneath them — they stop being tools and become a system that gets better every time it runs.

Section 2

LLM as Processor — the Core of Everything

This is the core of everything above — and there is only one idea to hold: the LLM is one processor, and it never changes. What changes is who drives it. First you do, by hand, adding more context and more tools as you climb the altitudes. Then the real unlock: you take the same pieces you built — context, skills, tools — and stand them up in a harness so the processor runs without you, always on. That single shift, from a person operating it to a system operating it, is what turns a chat into an agent.

Where you engage it — the human-operated altitudes (1–6)

Every altitude here is a person operating the processor. When you open ChatGPT or Claude you are already using it — inside a chat someone else built for you. Climbing is nothing more exotic than adding context and tools: you can chat with it, connect it to your data, automate it, code with it, and build whole systems and interfaces on it. Same discipline the whole way up — only the size of the context and the reach of the tools grow.

AltitudeWhat you're doingSurfaceWhat it adds
1 · Web chat + curated contextChatclaude.ai · ChatGPTA single conversation on a hand-assembled context — their chat, their model. Just talking to the processor, and it carries you further than most people realize.
2 · Chat + connectorsConnectProjects · MCP · Live ArtifactsPersistent files plus connectors that reach into Supabase / Drive / GitHub — the first real knowledge base and the first tools hung on the chat.
3 · Desktop + CoworkAutomatelocal files · scheduledLocal file access and an autonomous co-work mode, plus scheduled and phone-dispatched runs — the chat starts doing work while you are away.
4 · Claude CodeCodeCLI · cloudRepo-level instructions, reusable skills, subagents, and triggered runs with the laptop closed. Where we build the reusable pieces day to day.
5 · IDE & CascadeBuild systemsdeep repo + DBA full editor with deep repo context, multi-file edits, and a database behind it. The context is now an entire system and you are shipping software, not chatting.
6 · Full harnessStand upsovereignThe context and tools are wired into a system that runs on its own. No longer a chat — this is the bridge into agentic, below.
Altitudes 1–6 are all user-operated — a person is still in the loop. Start at the lowest altitude that fits the job; climb only when the context outgrows the room. Altitude 6 is where it stops being a chat and starts running on its own. The full teaching version lives on the Marketing OS ↗
The hinge: the pieces do not change across the line — context, skills, and tools are identical whether you are driving or a harness is. Only the driver changes: a person, or a harness that never sleeps.

How it runs without you — standing up always-on agents

This is the real power. Take the same well-designed pieces you just built and wrap them in a harness, and the processor executes on a trigger, a schedule, or an event — no one typing. Standing them up so they are always on is the whole game. There are three basic shapes, simplest first — a "real agent" is not better than a one-shot, just more expensive and less predictable, so use the simplest one that closes the task. Think of these as the primer; §6 builds them for real (the eight-slot anatomy + the harness menu).

PatternWhat it isLive at BC
One-shotSingle pass in, structured output out — and it can fire on a trigger, no human in the loopClassification Edge Functions · a node in a BC Pipeline · structured brain_chat
LoopRuns itself — iterates, or strings steps together, until the work is doneGrind (verify-loop, tests are the checker) · director task execution
AgentReasons over its tools, chooses its next step, holds stateCortex Conductor (supervisor) · the directors
Compose many of these into one system and you get a fleet. A fleet is the top of the scale — agents running in harnesses, coordinated, reading from the same shared context. That fleet is the agentic operating system this whole page is about, and what we are building toward (80 agents live today). See Harnesses & Fleet (§6) and the Agents Library. And the first, load-bearing piece you build for any fleet is the shared context every agent reads from — the Modular Contextual Spine, next.
Section 3

The Modular Contextual Spine — the System's Shared Context

Now that the processor is established, this is the other key element: the software we build around it is the spine — the durable, shared context every level of the build reads from. It holds who we are, how we look, what we say, who we sell to, and how the system actually runs, and everything after this section (skills, tools, agents) resolves its context here. It lives in two places: author and refine in git, serve structured in SQL — two substrates, each good at different things. (This is context, not memory — memory is a separate, emerging layer covered in §7.)

Substrate A · refine
Git — the authoring truth
A Karpathy / Wikipedia-style navigable knowledge graph: interlinked markdown with deliberate jumps that humans refine on the backend and that every agent and chat harness reads off. Diffable, reviewable, replayable, agent-native. Holds methodology, skills, rules, workflows, brand bibles, positioning, canon docs. This is how we refine the whole system.
Fails at: concurrent writes corrupt silently · grep degrades on paraphrase · no telemetry
Substrate B
SQL — the serving truth
Concurrency, row-level access, per-tenant scoping, semantic search, per-invocation analytics. Holds operational state, embeddings, CRM rows, resolved brand tokens served to apps, run logs.
Fails at: opaque without tooling · awkward to review · drifts from authored truth
Why this is commercial, not academic. Because brand bibles, ICP rules, and positioning are authored as structured files and projected into per-tenant SQL, a downstream skill can be contextually engineered for any tenant without being rewritten. That projection is what makes 87 skills templateable instead of BioCreative-specific.
Brand Bible
How we look — tokens, type, logo, layout
+
Content Spine
Who we are — offering, segments, personas, messaging
→
Skills & Agents
Plug into the spine — never brand- or context-hardcoded
→
Channels
Deployment — email, LinkedIn, SEO, print, website (not yet built)

The two halves of the spine

Deep page · real data
Brand Bible
BioCreative's actual brand rendered from the tokens: palette, type scale, spacing, contrast, logo, mascot. The visual half.
Resolved from design-tokens.json + brand.json
Open the Brand Bible →
Deep page · real data
Content Spine
Who we are, what we sell + price, who we sell to (segments + BCAS), personas, and messaging angles. The commercial half.
Resolved from the spine files
Open the Content Spine →
Tool
Brand Spine Matrix
The multi-tenant scoreboard: every client pack graded against the standard — required files, tier, WCAG contrast, owned-vs-inherited tokens.
7 client packs · scaffold: yes
Open the matrix →

Commercial context — resolved once, used everywhere

The keystone
Contextual Spine
The single index over identity, offering, segmentation, messaging, personas, proof, and funnel — canonical lines locked, every source file mapped.
Open the spine →
Segmentation
ICP Segments & BCAS
The 5-tier service-provider universe (Segments A–M) + the one canonical scorer, BCAS v2.1 — the deepest artifact in the spine.
Open segments →
The proof
Wiring & Traceability Matrix
Per artifact — SOW, proposal, deck, email, website, agents — which spine fields it consumes and whether it is wired, partial, or drifting.
3 wired · 7 partial · 3 drift
Open the matrix →

The system, instantiated — process traces & data

The spine is not only documents; it is the lineage of how a record moves through the system. These are the shipped end-to-end traces.

Diagram
Account & contact lifecycle
How a record enters, gets classified, scored, enriched, and routed.
Open →
Diagram
Commercial pipeline
The revenue path from signal to signed.
Open →
Diagram
Data inputs map
Every source feeding the system.
Open →
Trace
End-to-end traces
The W1–W7 paths: how a signal becomes a scored account, a message, a meeting, a deliverable.
Read →
Projection without validation is a slower way to be wrong. On 2026-08-03 we found CellScale had never once rendered in its own brand — the resolver silently dropped its nested colour objects and fell back to BioCreative teal for two months. Nothing errored, because silent inheritance was documented behaviour. The fix was to make a client pack with no primary colour fail loudly.
Reference
File vs DB — the field evidence
Interface-vs-substrate framing, when each wins, why hybrid is the mainstream answer, and the polyglot-persistence anti-pattern.
Read the analysis →
Reference
Context ≠ Memory
Why the spine (context / knowledge, this section) is the read substrate, and memory is a separate write layer — the learning-events path covered in §7. Keeping the two straight is what stops a system from hoarding context but never getting smarter.
Read the ladder →
Continually refined — and memory is not the spine. This context is kept current the way real systems are: backend refinement through git (the primary path today — humans editing the knowledge graph, projected into SQL). The second, emerging path is memory — learning events gathered as the fleet operates that eventually rewrite context at the graph level. That belongs to Retrieval & Graph (§7), and we deliberately keep it separate from the spine.
Section 4

Skills — the Composition Ladder

Capabilities compose upward: atom → bundle → workflow → agent → fleet. These are counted registry facts, not metaphors. Every skill carries its own classification in front-matter per the classification standard.

220
skills total
6
departments
108
atoms
6
bundles
106
composite use-cases
17
capability domains
25
client deliverable
87
templateable
8
out of scope
93
L1 starter
84
L4 workflow
The number that matters is 87. Templateable means one step away — a tenant binding, not a rewrite. Turning that into revenue is a binding problem (tenant spine + tools + credentials), not a build problem. Only 25 are deliverable today.
Top capability domainsSkills
research34
outreach29
sales22
infra20
document17
design16
media-video12
orchestration10
Browse
Skills Library
All 220 skills, filterable by disposition, level, composition, capability, and invocation. Shows the composition graph, the tools each skill uses, and which agents load it.
Generated from SKILLS_INDEX.md + front-matter
Open the library →
The contract
Skill Classification Standard
The front-matter contract: disposition, composition, level, capability (16 closed domains), handoff blockers, spine inputs, and the tools field.
Read the standard →
Section 5

Tools — the Fourth Piece

A skill says how. A tool is what it acts on: an external surface with an API, a credential, and usually a bill. Until 2026-08-03 there was no tool registry at all — the field existed on the skill schema and was empty on every skill.

83
tools registered
217
skills declaring tools
Backfill pending. 217 of 220 skills declare their tools. Until that reaches 100%, the “which skills break if this tool goes away” query is incomplete.
The honest framing, and it should stay consistent everywhere: the vendor tools are commodity subscriptions. The edge is the custom interconnection we built between them — staging pipelines, dedup, classification, message generation, reply routing, the modular context. The cost is real regardless.
Browse
Tools Library
Every tool by class, category, access surface, auth model, cost shape, host, and client-portability — each joined to the skills and agents that use it.
Generated from TOOLS_INDEX.md
Open the library →
The registry
TOOLS_INDEX.md
The authored source: controlled vocabulary, one row per tool, immutable slugs. Lint-enforced against every skill's tools field.
Read the registry →
The money
Tech Stack & Cost Buckets
Four buckets — find & enrich, hosting, outbound, build tooling — with how they interconnect and how each maps to a pricing tier.
The authority on cost
Open the buckets →
The load-bearing column is portable. A tool marked no_bc_oauth works only under BioCreative's OAuth grant, so any skill depending on it can never be client-deliverable without re-architecting per-tenant auth. That constraint — not build effort — is what caps the templateable count.

Tech stack & infrastructure — where an agent lives

A tool is an external surface; the stack is where our own agents run, are hosted, and hold a domain. Everything above plugs into this fabric.

HostRoleNotes
bc-opsHeart — heart-api, n8n, render services, ~50 containers, all static sitesTraefik on root_default. Claude Max — the default target for delegated work
bc-kbBrain — brain-api, LiteLLM, Ollama, the director ring, Cognee+Neo4jMetered API key — opt-in only
bc-made · bc-cubicbio · bc-spancorrClient tenantsNever used for BC processing
Database
Hub — operational
CRM, campaigns, pipeline, Edge Functions, the operational spine.
mjsgtszehjltxmbxtctz
Database
Brain — reasoning
Embeddings, agent memory, council decisions, research.
sfursphaicrvazkafbdi
Per tenant
Client spokes
One database per client, hub-and-spoke sync. SpanCorr on WIIZ, CARR, Made Transfer.
One workspace per client — hard rule
n8n's actual role: the general-purpose workflow layer for automation that does not need reasoning — polling, routing, notifications, glue. It runs on the VPS, self-hosted, across three tenants. It is not the agent runtime; that is the harness layer in section 6.
Section 6

Harnesses & the Fleet

Once you know the context, skill, tools, and shape, the harness is the environment that runs it: process model, state, retries, credentials, observability, cost profile. We keep a build-agnostic menu — patterns you can build on, never BC's own products.

The anatomy of one agent

Every agent is the same processor with eight engineered slots filled in around it — context, skills, tools, and memory from §1, plus shape, breakers, trigger, and eval — all inside a harness (the frame) that runs it and writes back to the spine. Build one by picking the shape, then filling the slots. A fleet is many of these, coordinated.

Agent anatomyThe LLM is a processor. Everything we engineer around it is the product.HOST · a container on a VPS — the box the whole thing runs inCONTEXTwhat it must knowMCS slicebrand spineRAG read-pathSKILLShow to do itmethodrulesworked patternsTOOLSwhat it acts onvendor APIsBC servicesMCPMEMORYwhat it keepstypes heldwrite-pathstoreSHAPEdepth × widthloop = depthswarm = widthBREAKERSwhen it stopsmax_turns$ capwall-clockdup-callTRIGGERhow it's invokedscheduledeventappagenton-demandEVALhow we knowheartbeatrun-logsLLMPROCESSORswappable · via LiteLLMWRITES BACKNew skills, corrected facts, proposed memory, updated registries return to the same modular contextsystem. A system that consumes context but never returns to it degrades; one that writes back compounds.
27
harness patterns
5
VPS
2
core databases
NeedHarness
One-shot structured outputPydantic AI single agent
Multi-step + tools + statePydantic AI in the brain-api pattern — the default
Long autonomous coding (>15 min)Claude Code on bc-ops via /delegate-vps
Quality gate (verify-loop)Grind / Claude Agent SDK
Genuine parallel widthDynamic Workflows / Antigravity
Deterministic DAGBC Pipelines on the Archon engine
Heartbeat / continuousWorkflow Console
The Karpathy principle governs this menu. "Six months is an eternity in this space." Default to strong, well-supported primitives and refine our own method rather than chase frameworks. Every new harness gets a 30-day skeptical window. We are building an iOS-style consistent system, not a museum of frameworks.
Browse
Methodology Map
The narrative plus the full harness menu with production status, the loop/swarm vocabulary, the intake flow, and which live agents run on which harness.
Open the map →
Browse
Agents Library
Every deployed agent and service with its harness, skills, trigger, host, and health.
Open the library →
Naming
Ours vs attributed
Coined external names stay attributed in the learning layer (Archon, Ralph, Dark Factory). What we build gets a BC name: BC Pipelines, Grind, MCS.
Read the glossary →
Registry
Agent Fleet Registry
The living deployment source of truth — every container × VPS with wiring and health.
Open the registry →
Registry
BC Fleet Matrix
The correlation surface: for agent X, the whole picture — harness, skills, trigger, host, usage.
Open the matrix →
Section 7

Retrieval & Graph — the Honest State

This is the least-built part of the methodology and is described as such. The 2026 consensus is that retrieval is a three-stage pipeline — retrieve (BM25 + dense) → fuse (RRF) → rerank. BioCreative currently runs none of it.

Works
Query-time retrieval Live
brain-search, brain-router, brain-recall, meeting_search. Pennies per month. /brain-query, the directors' query_brain, and the Conductor all function.
91,717 Hub + 194,932 Brain vectors retained
Frozen
All auto-embedding Frozen
Since 2026-08-03. July's Google bill was $416, of which $404 was 2.02B embedding tokens — a pagination bug re-embedded ~96% of the corpus 4×/day while logging $0.0000 from a hardcoded zero.
Six embedders now default off behind EMBED_ENABLED
Missing
Hybrid, rerank, graph-in-path Not built
No BM25, no fusion, no reranking. Cognee/Neo4j runs and syncs but no production agent traverses it. The literature's single highest-leverage component — reranking — is absent.
Found 2026-08-03: the RAG quality CI (.github/workflows/rag-quality.yml) triggers on scripts/embed_knowledge_base.py, but the embedder moved to departments/internal-ops/scripts/ in the August reorg. That trigger is dead — embedder changes no longer run the eval. A dead quality gate is worse than none, because it implies coverage we do not have.
The cheapest meaningful improvement is not a graph and not a better embedding model. It is adding BM25 + RRF + a reranker to a corpus we have already paid to embed. Benchmarks: hybrid + rerank reaches Recall@5 0.816 vs 0.695 for hybrid alone and 0.587 for dense-only. None of it requires turning embedding back on.
Research
Graph & Retrieval Engineering
The field evidence (GraphRAG vs vector, hybrid + RRF + rerank, entity resolution as the real precondition), plus BC's unflattering audit and a cheapest-first path.
Read the reference →

Memory — the emerging write-back layer

Retrieval and the spine (§3) are the read path — what the system knows. Memory is the write-back path, and it belongs here, not in the spine. As the fleet operates it produces learning events — corrected facts, resolved entities, outcomes, new skills — that should periodically roll back up to update the graph and, eventually, rewrite context at the top. This is distinct from git/backend refinement (the primary path today): memory is the agentic second path, and it is the least-built idea on this page.

Honest state: today, context is refined by hand through git and projected into SQL. A true memory loop — learning events aggregated and promoted into the graph — is designed, not built. When it exists it will run propose-then-promote: a run proposes, a human or gated check confirms, and only then does it become canonical. No agent silently rewrites production truth.
Section 8

UX & Operator Surfaces — the Eyes

Everything above is the machinery: data organized, software written, agents built, context stored, work executed. None of it is operational until a human can see it and steer it without living in a terminal. That layer is the UX — the Eyes of the four-part system introduced in section 2. It is where the whole build stops being infrastructure and becomes something a non-engineer runs.

Brain
Knows — the Contextual Spine + retrieval
+
Heart
Circulates — heart-api, n8n, schedules
+
Nervous system
Executes — skills, tools, the fleet
→
Eyes (this section)
Steers — the UX a human operates from

The operator surfaces we run

OS surface
Internal OS
This intel hub — the generated launcher every registry renders into. The operator surface for how we build.
You are here
OS surface
Marketing OS
The prospect-facing story surface — tiers, ladder, anatomy, the org map. Same method, outward-facing.
Open ↗
Client front-offices
CARR · Made · SpanCorr · Launch
Per-client React/Vite dashboards (Lovable-built, Supabase-backed) where each client sees and steers their own commercial engine. One front-office per tenant.
Separate repos, one per client
Internal dashboards
Cortex Inbox · Account Atlas
The propose-then-confirm surfaces: Cortex approvals, account/contact atlas, campaign dashboards. Where an operator makes the call a human must make.
In-app on the Hub
Libraries
Skills & Agents
The browsable registries — every skill and every deployed agent, filterable, cross-linked. The operator's map of the fleet.
Open →
Emerging pattern
MCP-UI & Live Artifacts
Answers that render as dashboards inside the chat and refresh when reopened — the interface layer collapsing into the conversation. Tracked, not yet a product.
Watching
This is the least-standardized layer we have — each surface was built for its moment, and unifying them (shared design system, shared auth, shared component kit) is live work. Every surface resolves the same brand tokens, so they at least look like one system. The full inventory, per-surface, is the deep page.
Deep page
UX & Operator Surfaces
The full operator-surface inventory — every OS surface, client front-office, dashboard, and emerging pattern, with its stack, host, and the Eyes role in the four-part system.
Real inventory
Open the inventory →
The design tie
Brand Bible
Every operator surface renders the same tokens — the reason they read as one system even though they are separate apps.
Open the Brand Bible →
BioCreative OS surfacesInternal OS · this intel hubMarketing OS · the prospect story ↗Org Map · the agent org chart ↗