Data audit & quality assessment
What data you hold, where it lives, how good it is, who owns it and what is blocking you. Delivered as a map plus a prioritised fix list with effort estimates, not a lecture on best practice.
Service · Foundations
Nobody puts “data plumbing” in a board paper. But almost every stalled AI project we are called in to rescue stalled here — duplicated customers, numbers that disagree, and a critical spreadsheet on one person’s laptop.
What is AI readiness?
AI readiness is the degree to which an organisation’s data, systems, processes and governance can support useful artificial intelligence. In practice it means: the data exists, it is accessible through something better than a manual export, it is accurate and de-duplicated enough to be trusted, definitions are agreed across departments, access is controlled, and there is a lawful basis and retention rule for the personal data involved.
There is a version of this work that costs a fortune and takes two years, and we are not proposing it. Enterprise data transformation programmes have a poor record everywhere, and a worse one in mid-market Kenyan businesses where nobody has a spare team to run them.
Instead we fix the data needed for the next thing you want to build, and only that. If the first automation needs clean supplier records, we fix supplier records. Customer records can wait until they block something. This sequencing is the difference between a data project that delivers and one that becomes a permanent line item.
Signs you need this
In practice
We recommend starting with the audit — it is cheap, fast, and it usually changes what you thought you needed.
What data you hold, where it lives, how good it is, who owns it and what is blocking you. Delivered as a map plus a prioritised fix list with effort estimates, not a lecture on best practice.
Scheduled, monitored, retrying flows that move data between your systems reliably — replacing the manual export-and-import ritual that fails whenever someone is on leave.
Agreed definitions, identity resolution across systems so one customer is one customer, and one place everyone queries. Half technical, half diplomatic.
The daily, weekly and monthly numbers assembled automatically and delivered where people already look — including WhatsApp, which in Kenya beats any portal for actually being read.
Who can see what, retention schedules, audit logging and the documentation the Office of the Data Protection Commissioner would expect you to produce.
Getting historical records out of filing cabinets, old systems and dead spreadsheets into something queryable — often using document AI, which makes this far cheaper than it used to be.
The same answer, two ways
What “fixing the data” actually means
Getting your business to agree on the facts, and keeping them in one place.
Most companies do not have one set of numbers. They have several, and each department trusts its own. Sales counts a sale when the order is signed; finance counts it when the money lands; the store counts it when the goods leave. None of them is wrong, but nobody has written the differences down, so every management meeting starts with an argument about whose figure is right.
The first job is agreement. We get the relevant people in a room and write down what each number means and when it counts. This is unglamorous and occasionally tense, and it is the highest-value hour in the whole project.
The second job is joining things up. Right now your customer might be “Kamau Enterprises” in one system, “Kamau Ent Ltd” in another and a phone number in a third. We match them so that one customer is one customer, and you can finally see what any of them is actually worth to you.
The third job is making it automatic. No more Friday exports. The numbers assemble themselves, on schedule, and arrive where you already look — a dashboard, an email, or a WhatsApp message at 07:00.
Once that exists, AI becomes straightforward. Most of what looks like an AI problem was a “we could not see it” problem all along.
Pragmatic ELT into a modelled store, with identity resolution, tested transformations and observability — sized to the organisation, not to a reference architecture.
Assessment. Source inventory with extraction feasibility per system, profiling for null rates, cardinality, referential integrity and duplication, plus a lineage sketch for the numbers that appear in management reporting. Personal-data mapping runs in the same pass, because it needs the same inventory.
Ingestion. Change-data-capture or incremental extraction with watermarks where the source permits; full snapshots where it does not. Landed raw and immutable first, so a transformation bug never means re-extracting from a fragile production system. Orchestrated with retries, alerting and backfill capability.
Modelling. Layered — raw, staging, marts — with transformations in version control and tested (uniqueness, not-null, referential, accepted values, freshness). For most Kenyan mid-market organisations this lives comfortably in Postgres; BigQuery or Snowflake when volume or concurrency genuinely justifies the cost.
Identity resolution. Deterministic matching on strong keys — KRA PIN, national ID, phone number in normalised E.164 — then probabilistic matching on name and address with a human review queue for the uncertain band. Kenyan name variation, initialisation and transliteration means blind fuzzy matching produces false merges, which are far more damaging than missed ones.
Semantics. A metric layer where each business definition is written once and referenced everywhere, so “active customer” cannot quietly mean three things.
Governance. Role-based access enforced at the data layer, column-level masking for sensitive fields, retention policies implemented as scheduled jobs rather than intentions, and audit logging of access to personal data — the evidence you need for a Data Protection Act, 2019 enquiry.
What you get
Foundations that make every later AI project cheaper, faster and less risky.
Typical engagement
Every quote is fixed before work starts. If we scoped it wrong, that is our risk, not yours. See how pricing works.
Typical stack
We choose tools your team can maintain, not the ones that make us look clever.
Questions
Because an AI system is only as good as what it can see. If your customer exists three times under slightly different names, no model can tell you how much that customer is worth. If your stock figures are wrong, a demand forecast built on them will be confidently wrong too. Most AI projects that quietly fail did not fail on the model; they failed because the data underneath was fragmented, duplicated or simply absent — and nobody checked before the budget was approved.
Probably about the same as everyone else’s, which is worse than you think and less catastrophic than you fear. The common pattern we find in Kenyan organisations: a reliable transactional core, a duplicated and inconsistent customer master, spreadsheets holding critical logic that exists nowhere else, and one system whose data nobody trusts but everyone uses. A two-week audit tells you exactly where you stand, with a prioritised fix list.
Often not, and we will say so. If you have three systems and a few million rows, a well-designed Postgres database with scheduled syncs will serve you for years at a fraction of the cost and complexity. A warehouse earns its keep when you have many sources, real volume, or several teams querying independently. Selling infrastructure a Kenyan mid-market business does not need is one of the more expensive things a consultancy can do to you.
An agreed answer to questions like “how many customers do we have?” or “what did we sell last month?” that everyone in the business uses. Today those numbers usually differ by department because each pulls from a different system with different rules and different cut-off times. Establishing a single source of truth is partly technical — pipelines, definitions, identity matching — and partly political, because someone has to accept that their familiar number was wrong.
Substantially. You cannot comply with the Kenya Data Protection Act, 2019 if you do not know what personal data you hold, where it lives, who can see it and how long you keep it. The data mapping we produce for AI readiness is the same artefact your compliance obligations require, so the two pieces of work should be done together rather than twice.
Yes, and we usually do, because a dashboard is how the data work becomes visible to the people who paid for it. We build in Power BI, Looker Studio or a lightweight custom front end depending on what your team can maintain. The rule we hold to: a dashboard must answer a specific recurring question that someone actually asks, or it will be opened twice and abandoned.
Keep exploring
Most clients combine two or three of these.
We find the repetitive work eating your team's week — reconciliation, data entry, approvals, reporting — and hand it to software that never gets tired.
A WhatsApp and website assistant that answers customers in English and Kiswahili, qualifies leads, checks order status and escalates to a human when it should.
An honest map of where AI will make you money, where it will waste your money, and what to do first — costed, sequenced and tied to your numbers.
Start here
Book a free 30-minute call. We will map one process end to end, tell you honestly whether AI is the right answer, and put a number on what fixing it is worth. No slide deck, no jargon.
Nairobi-based · We reply the same working day · English & Kiswahili
Talk to a human