Current work / field notes

AI, applied to the operating system of a business.

BlackBoxx.ai is where I build, test, and pressure-test practical AI systems—connecting agents, applications, people, controls, and decisions.

I am a 25+ year founder-operator and CFO. I have scaled a company from zero to $30M+ ARR and more than $200M in cumulative revenue, and built finance functions from the ground up. These tools get evaluated against how finance actually gets done—close, forecast, board pack, audit—not against a demo script. Karl Ohlemann ↗

Independent AI Systems LabThe work ranges from live operating infrastructure to early concepts and product evaluations. Each is labeled accordingly.

01Inputs 02Specialized agents 03Validation + exceptions 04Executive decision

In the lab

Building systems that do useful work—not AI theater.

The focus is operating leverage: giving capable people better infrastructure, shortening the distance between information and action, and keeping a human accountable for the consequential call.

Active system / Agent operations

A hierarchical 16-agent operating environment

I run 16 specialized agents across 15 namespaces—research, finance and analysis, technology operations, executive support, legal review, CRM, and chief-of-staff coordination—with an orchestrator acting as the control plane. It is a governed fleet with role separation, not a pool of isolated chat sessions. Two runtimes operate side by side today, Claude Code and Codex; local inference through OpenCode and Ollama is planned once runtime routing is defined.

Most of the engineering is not prompting. It is authority. Agent-to-agent communication started as open collaboration and was rebuilt as explicit allow-listing: no wildcard routes, 34 directional rules, one-way where a domain should receive work but never initiate it. Every one of the 15 non-orchestrator agents runs under a capability ceiling that blocks six destructive capability classes. Denied routes were live-tested rather than assumed, because a control you have not tried to break is a control you have not verified.

The principle underneath all of it: agent intelligence and agent authority are separate things. An agent can be highly capable while specific destructive actions and cross-domain communication routes are technically denied. That distinction is the difference between an impressive demo and infrastructure you can begin to trust with real work.

Built on DorkOS, open source (MIT) ↗

Operating model / People + systems

AI-enabled teams, documented work, clearer accountability

I am working out how process engineering, documented SOPs, automation, and AI-assisted execution function as one managed operating system—where the work is defined well enough that a person or an agent can run it, and accountability for the outcome stays with a named owner either way.

The hard part is triage, not tooling. Most work sorts into three buckets: repeatable and rule-bound, where an agent is a genuine gain; judgment-bound but documentable, where a person runs it against a written standard and an agent assists; and consequential or ambiguous, which stays with the operator and should never be quietly automated. Getting that sort wrong in either direction is expensive—automate judgment and you inherit silent errors, hand routine work to people and you pay for capacity you did not need.

So the artifacts matter more than the models: a written procedure, a defined owner, a checkpoint where someone confirms the output is right, and a record of what happened. That is the same discipline a finance function runs on. It is what makes AI-assisted work reviewable instead of merely fast.

Under evaluation

AI-native finance, tested through a CFO’s lens.

I am demoing emerging finance platforms to understand where AI creates genuine operating leverage—and where trusted data, controls, traceability, and human judgment still determine whether the output is useful. The notes below are preliminary observations from product demonstrations and early evaluation—not implementations, and not endorsements. Each platform is assessed against the work finance teams actually perform: close, forecast, reporting, controls, communication, and decision support.

What I am testing for

  • Data integrityWhether figures reconcile to a governed source rather than a model's recollection.
  • AuditabilityWhether an answer can be traced back to the transaction that produced it.
  • PermissionsWho can see and change what, and whether that holds up under review.
  • Workflow fitWhether it matches how the close, the forecast and the board pack actually run.
  • Scenario speedHow fast a real driver change produces a defensible re-forecast.
  • Human reviewWhere judgment is required, and whether the system makes that explicit.
  • Integration depthWhether it reaches the ledger, the bank and the field—or only the spreadsheet.
  • Decision usefulnessWhether an executive would act on the output without rebuilding it first.

Construction finance

Adaptive.Build

Purpose-built for construction, and it goes well beyond the FP&A boundary where most AI finance platforms stop. Agents receive, read, and code invoices to specific jobs and cost codes, route approvals, support billings, maintain real-time WIP visibility, and assist with draws and lien-waiver workflows. The distinction that matters: AI applied to the repetitive transaction-level work connecting the field, project accounting, and the books—not only to analysis. Not to be confused with Workday Adaptive Planning.

adaptive.build ↗

AI-native FP&A

Drivetrain

Feels genuinely AI-native rather than an established planning platform with an assistant bolted on. Model generation, plain-English data transformation, scenario planning, anomaly detection, and traceability inside an unusually nimble environment. What stood out in the demo was both ease of use and the development model behind it—tailored product updates reportedly delivered in days rather than the heavy implementation cycles typical of enterprise FP&A.

drivetrain.ai ↗

Spreadsheet-native FP&A

Aleph

A robust FP&A platform built around the Excel and Google Sheets environments finance teams already use. No-code connectors and bidirectional spreadsheet integration connect existing models to a centralized, governed data layer without asking anyone to abandon familiar workflows. Agents make analysis, reporting, and ad hoc scenario work easier to execute, and scheduled reports with automated distribution carry the work through to stakeholders.

getaleph.com ↗

Excel-native FinanceOS

Datarails

The most compelling advantage is breadth. It consolidates data from more than 600 systems while preserving a team's existing models and formulas, then adds centralized version control, cell-level audit trails, automated executive dashboards, and governed data lineage. Expanding modules reach past conventional FP&A into close management, cash visibility, spend control, and AI-assisted strategy—a broader finance operating layer that does not require giving up Excel.

datarails.com ↗

The operating thesis

Intelligence matters when it changes the work.

The objective is not maximum automation. It is a better-designed system: routine work moves faster, exceptions surface earlier, decisions arrive with better context, and accountability remains clear.

01

Compress repetitive operating cycles.

02

Turn scattered information into usable context.

03

Escalate uncertainty and exceptions visibly.

04

Keep consequential decisions human-owned.

BlackBoxx.ai

Building, testing, and learning where AI meets real operations.

This is an active lab, not a finished catalog. The systems and tools will continue to evolve as the work produces better evidence.