Dogfooding · self-assessment record

We measure our own system
with the same method

KNDLI applies the same measurement and improvement method it proposes to clients to its own AI agent knowledge management system. Self-assessed on the public AKM Index v1.2 rubric, the score is 80.75/100, maturity band M3, up from 66.25 in round 1 to 72.75 in round 2 and 80.75 in round 3, a cumulative +14.5 in three weeks. It is an honest record, with five-pillar adversarial review and evidence gates to prevent inflation.

The surest way to show a method is more than words is to measure ourselves with it first. Below are the scores and the evidence.

80.75
total / 100
M3
maturity band
+14.5
self-improvement over 3 rounds
01

What does our
system score?

The public AKM Index v1.2 rubric scores a knowledge management system across five areas. Total 80.75, maturity band M3 (orchestration). We show the strong and weak spots as they are, nothing hidden.

AKM Index v1.2 · round 3

Prompt designManaging the instructions that give AI its work, systematically 16 / 20
Context designPreparing the materials and memory the AI can draw on 18.75 / 25
Harness designKeeping the environment where AI actually uses tools and does work safe 18 / 20
Operating loop designA repeating structure that runs, checks and improves every day 16 / 20
Interoperability · governanceLetting multiple AIs, devices and people work by the same rules, with safety mechanisms in place 12 / 15
02
SELF-IMPROVEMENT

How much can
the score rise?

Measuring once is not the end. We found the weak spots in round 1, fixed their causes and re-measured in round 2, then fixed more and re-measured in round 3. The gains come from adding verification loops and safety mechanisms, not from showing off new features.

Round 1
66.25

First measurement. Identified weak areas and gaps in evidence.

+6.5 self-improvement
Round 2 · 3 days later
72.75

Fixed the causes of silent failures and added a verification loop.

+8.0 self-improvement
Round 3 · 2026-08-11
80.75

Re-measured after routing tests, 100% metadata backfill and stronger governance. Cumulative +14.5 over round 1.

03

Can a self-assessed
score be trusted?

Scores you give yourself inflate easily. So we use three mechanisms to make this verifiable measurement, not self-grading.

01 · Public rubric

Public rubric

We measure with a public standard anyone can read (AKM Index v1.2), not a yardstick of our own making. The same table lets you compare with other organizations.

02 · Adversarial review

Adversarial review panel

Each of the five areas gets a review panel whose job is to push back hard. Its role is to find flaws, not to be generous.

03 · Evidence gate

Evidence gate

No supporting artifact, no points: the deduction is mandatory. A level claimed in words does not count; the score only rises with real records.

04
THE EVIDENCE

A record of the system fixing the system

These are the three level-4 artifacts the score rests on. All three actually run every day, and they are what lets the system find and fix its own problems.

Level 4

Persistent memory

What is learned once is remembered in the next session. No need to explain from scratch every time, so knowledge accumulates instead of scattering.

Level 4

Observability

What ran, when and how is kept on record. Even jobs that failed quietly show up in the logs, so the cause can be traced.

Level 4

Periodic review

The system re-checks and improves itself on a set cycle. This self-assessment is itself part of that review loop, and the next re-measurement is already scheduled.

05

So what actually happened

Behind an abstract score is a concrete day. Here is real work, shown as it happened, to see how this method runs in practice.

Found and fixed the causes of 3 silent failures in a single day

Observability logs uncovered three automations that had quietly stopped without a single error. We pinpointed the real cause behind jobs that looked normal on the surface and fixed them the same day.

Designed the daily verification loop and automatic memory capture ourselves

We built a loop that verifies each day's results on its own and made sure what was learned that day lands in team memory automatically. Even if a person forgets, the system remembers.

Guardrail hooks pause hard-to-reverse actions until a person approves

We put guardrails before hard-to-reverse steps like sending or payment, so even automated work must pass a human check. The design puts safety ahead of speed.

NEXT STEP

Apply this method
to your company's system

The measurement and improvement we used on ourselves is exactly what we apply to client AI systems. We start by diagnosing together where you stand now and what to fix first.

Ask for a diagnosis See AI automation consulting

※ The scores and story on this page come from measuring KNDLI's own system against a public rubric. Client materials and consultation details are never disclosed.

FAQ: knowledge management self-assessment

What is the AKM Index?

The AKM Index is a public rubric (v1.2) that scores the maturity of an AI agent knowledge management system from 0 to 100 across five areas: prompt design, context design, harness design, operating loop design, and interoperability and governance. Each area is leveled by observable behavior, and every point requires real evidence. KNDLI measured its own system against this public standard and scored 80.75, maturity band M3.

Why does KNDLI measure its own system publicly?

To show that the measurement and improvement method we propose to clients is the same one we apply to our own system. Instead of saying things will get better, we publish, against a public standard, what we scored, where we are weak, and how much we improved in how many days, with numbers and evidence. Measuring yourself first like this is usually called dogfooding, and it is the basis for trusting that the method actually works rather than being just words.

How is the score verified and inflation prevented?

Each of the five areas has an adversarial review panel, and any item without evidence is forced to lose points. To claim a level you must present artifacts that actually exist, such as persistent memory, observability logs and periodic review records. Because the score cannot rise without evidence, the structure itself prevents inflated results. The rise from 66.25 in round 1 to 80.75 in round 3 did not come from showing off new features; it came from fixing the causes of silent failures and actually adding verification loops and safety mechanisms.

Explore More Services
Full Portfolio AI automation consulting AI Visibility AI Automation Custom AI Training Profile Contact