Dogfooding · Self-Assessment Record

We measure our own system
with the same methodology

KNDLI applies the same measurement and improvement methodology we propose to clients to our own AI agent knowledge management system. Self-assessed on the open AKM Index v1.1 rubric: total score 72.75/100, maturity band M3, improved +6.5 points in 3 days from a first-round 66.25. An honest record, guarded against inflation by 5-pillar adversarial review and evidence gates.

The surest way to show a methodology is more than talk is to measure ourselves with it first. Below are the scores and the evidence.

72.75
Total / 100
M3
Maturity band
+6.5
Self-improved in 3 days
01

Our system, in numbers

The open AKM Index v1.1 rubric scores a knowledge management system across five pillars. Total 72.75, maturity band M3 (Orchestration). We show the strong and the weak spots exactly as they are.

AKM Index v1.1 · Round 2

Prompt DesignSystematically managing the instructions that direct AI work 15 / 20
Context DesignKeeping the references and memory AI needs at hand 15 / 25
Harness DesignKeeping the environment where AI actually uses tools safe 17 / 20
Operating Loop DesignA repeating structure of daily running, checking, improving 16 / 20
Interoperability · GovernanceMultiple AIs, devices and people collaborating under shared rules, with safeguards 9.75 / 15
02
SELF-IMPROVEMENT

A score we raised ourselves in 3 days

One measurement isn't the end. Round 1 exposed the weak spots; we actually fixed their root causes, then measured again in round 2. Not by showing off new features, but by adding verification loops and safeguards.

Round 1
66.25

First assessment. Identified weak pillars and gaps in evidence.

+6.5 self-improved
Round 2 · 3 days later
72.75

Re-measured after root-cause fixes. Band M3 held, total score up.

03

Three things that make the score trustworthy

Self-graded scores inflate easily. So we built in three mechanisms that turn self-scoring into verifiable measurement.

01 · Public rubric

Public rubric

We don't grade with a yardstick of our own making, but with an open standard anyone can inspect (AKM Index v1.1). The same table lets you compare against other organizations.

02 · Adversarial review

Adversarial review panel

Each of the five pillars gets a review panel whose job is to push back hard. Their role is not to be kind but to find flaws.

03 · Evidence gate

Evidence gate

No supporting artifact, forced deduction. A level claimed in words alone doesn't count; the score only rises when real records exist.

04
THE EVIDENCE

A record of the system fixing the system

Three Level 4 artifacts back the score. All three actually run every day, and they are what lets the system find and fix its own problems.

Level 4

Persistent memory

What is learned once is remembered in the next session. No re-explaining from scratch each time; knowledge accumulates instead of scattering.

Level 4

Observability

What ran, when and how is kept on record. Even jobs that fail silently surface in the logs, so root causes can be traced.

Level 4

Periodic review

The system re-examines and improves itself on a fixed cycle. This self-assessment is itself part of that review loop, and the next re-measurement is already scheduled.

05

So what actually happened?

Behind the abstract score is a concrete day-to-day. Here is how this methodology runs in practice, taken straight from real work.

Diagnosed and fixed 3 silent failures in a single day

Three automations had quietly stopped without a single error message; the observability logs surfaced them. We pinpointed the real cause behind work that looked fine on the surface, and fixed it the same day.

Designed a daily verification loop and automatic memory capture

We built a loop that verifies each day's results and automatically writes the day's lessons into team memory. Even when people forget, the system remembers.

Guardrail hooks stop hard-to-undo actions ahead of human approval

Steps that are hard to reverse, like sending or payments, sit behind guardrails: even when automated, they must pass human confirmation. A design that puts safety before speed.

NEXT STEP

This methodology,
applied to your systems too

The same measurement and improvement we apply to ourselves, we apply to our clients' AI systems. We start by diagnosing together where you stand now and what to fix first.

Request a diagnosis See AI automation consulting

※ The scores and narrative on this page are the result of measuring KNDLI's own system against an open rubric. Client materials and consultation details are never disclosed.

FAQ — Knowledge Management Self-Assessment

What is the AKM Index?

The AKM Index is an open rubric (v1.1) that scores the maturity of an AI agent knowledge management system from 0 to 100 across five pillars: prompt design, context design, harness design, operating loop design, and interoperability & governance. Each pillar defines levels by observable behavior, and every score requires real supporting evidence. KNDLI measured its own system against this open standard and scored 72.75 total, maturity band M3.

Why does KNDLI measure its own system publicly?

To show that the measurement and improvement methodology we propose to clients is applied to our own system in exactly the same way. Instead of saying "it's getting better," we publish, against an open standard, the current score, the weak spots, and how much we improved in how many days — all in numbers and evidence. Measuring yourself first like this is commonly called dogfooding, and it is the basis for trusting that the methodology actually works rather than being talk.

How are the scores verified, and how is inflation prevented?

Each of the five pillars has an adversarial review panel, and any item without evidence is forcibly marked down. To claim a level, you must present artifacts that actually exist — persistent memory, observability logs, periodic review records. Without evidence the score cannot rise, which structurally prevents overstating. The climb from 66.25 in round 1 to 72.75 in round 2 came not from showing off new features, but from fixing the root causes of silent failures and actually adding verification loops and safeguards.

Explore other services
Full Portfolio AI Automation Consulting AI Visibility AI Automation Custom AI Training Profile Contact