We don’t have to admire a company’s organizational culture to benefit from its data practices. When I read the article Uber’s Journey Toward a Better Data Culture From First Principles, I thought it offered great perspective into a disciplined approach that can be used to fix data quality issues in any kind of entity, big or small.
While most businesses don’t need or have the resources to implement the same technological solutions used by a company that processes petabytes of data on a single day, the problems mentioned in the article are pervasive. They equally affect organizations in the public, private, and non-profit sectors, of any size, and whether or not they’re facing scalability issues.
Below I provide some real-life examples of the same issues mentioned in the article happening to small and mid-size organizations. If you find yourself nodding your head in acknowledgement, make sure to check the last part of the post for the actions you can take to overcome those problems.
1) Duplication, inconsistency, confusion, and extra work caused by the absence of a single source-of-truth
A mid-sized company had Sales Operation, Product and other departments calculating the same business measures using their own formulas. The need for complicated formulas resulted from a poorly structured customer database. For instance, if you weren’t careful, you might count the same customer twice, a mistake that would inflate metrics like conversion rate and annual recurring revenue (ARR). Since each group made their own calculations, the numbers never matched, and internal reports inspired little confidence. With conversion rate and ARR being such vital metrics for the business, one would think that having a single source of truth would quickly become a top priority. Still, while the company’s leadership team agreed that there was an issue, it never took the necessary steps to fix it.
2) Inefficiencies caused by lack of a data inventory
An analyst was asked to create a business report that required NPS (Net Promoter Score) ratings filtered by respondents who were the decision-makers at the time the survey was taken. The analyst talked to person A in the Customer Support team who referred them to person B in the same team. B had the data for the current year, but for past data, the analyst would have to talk to person C in Sales Operations… Collecting the data required a process that took days when it should have taken minutes.
3) Inefficiencies and lack of confidence in results caused by disconnected systems and tools that don’t integrate with each other
A marketer was writing a promotional email for prospects, and had to spend a lot of time trying to reconcile the lists of contacts grabbed from two systems. The situation not only created delays and inefficiencies, but caused real damage when a promotional email offering an attractive discount for new clients ended up in the inbox of a customer that had already signed a one-year contract at a higher price.
4) Erroneous findings caused by poorly instrumented logs
A company that offers an app to customers implemented a tool to log feature usage. Because of poor instrumentation, the logs provided an incomplete picture. The product team ended up making the wrong decision to prioritize bug fixes for feature A and delay the fixes for feature B based on the information that 65% of customers used feature A and only only 23% used feature B. In reality, 87% of customers used feature B, but due to logging constraints, only part of the activity in that feature could be tracked.
5) No quality guarantees and inconsistent response time for bug fixes caused by lack of data ownership & SLAs
Reusing the previous examples, imagine that your company has two competing needs: fix the logging system so the product team can have a full picture of user behavior in order to improve feature prioritization, and reconcile two lists of business contacts to ensure customers who opted out of promotional messages are excluded from a contact list. Both tasks require involvement from the same understaffed development team; what gets done first?
If the relative value to the business of accurate data in both domains is unclear, the answer may be decided based not on the importance of the data for business-critical decisions, but on other considerations such as which manager complains more loudly or has more access to an influential stakeholder to get their request placed at the top of the priority list.

How to overcome those issues even when you’re small and resource-constrained
You don’t need to have a billion-dollar business to be able to fix data issues. As pointed out in the article by Puttaswamy and Srinivas, the answer is to focus on first principles. Those are the fundamental truths from which we can reason about the best course of action not based on how things have always been done, or how others solved the same problem, but by breaking the situation down into its component parts and seeing what’s possible. That is the work of first-principles thinking, and here’s how to apply it to your data quality problem:
1) Identify the key data assets the business needs to run and the performance measures it needs to monitor, and direct your efforts to developing a single source of truth for them.
A key data asset for a business may include up-to-date information about opt-outs to avoid having messages accidentally sent to unsubscribed users being reported as spam. For another, it may be accurate parts inventory data to prevent costly issues with order fulfillment.
Likewise, in terms of performance measures, even two businesses of the same size in the same industry may need to focus on entirely different factors based on their strengths and weaknesses. It’s possible that for company ABC the customer satisfaction index drives a lot of important decisions, while for competitor XYZ customer satisfaction has remained stable for years and what really matters is sales/employee and the percent of sales from proprietary products.
Once you know what is important to track, it’s time to focus on establishing robust processes to calculate top-level measures in a centralized manner, with standard quality checks to ensure accuracy, timeliness, and easy consumption of relevant indicators by anyone who needs them. This may sound complex and expensive, but having scarce resources is no excuse for sticking to inefficient processes that end up costing more in manual labor and rework when staff has to keep validating and correcting the data or recalculating the same performance measures all the time. A couple of examples of businesses using low-budget solutions to achieve this goal can be found below.
2) Create an inventory of your important data assets, measures, and reports that are routinely produced.
No matter its size, any company can and should maintain a systematic inventory of the data assets on which people rely to make business-critical decisions. In a company that is heavily dependent on NPS score to gauge the “health” of its customer base, there is no excuse for an analyst having to spend days tracking down metadata about NPS surveys in order to produce a new business report.
Of course, you shouldn’t dedicate resources to documenting all available data assets. Only the data that is critical for the business as a whole or for one or more of its departments or functions must be cataloged.
The many benefits of maintaining a data inventory include:
identifying duplicative work (e.g., sales and marketing both calculating conversion rate) and measurements or reports that lost their utility (e.g., number of customers still using a legacy version of the product when the number is now one and no longer changing over time) that can be eliminated;
providing visibility into valuable data assets that can be given new, strategic uses;
eliminating inefficiencies caused by people who need a data asset having to track down where the data is, figure out in which format it’s stored, etc.;
preserving valuable contextual knowledge that end up lost when an employee leaves the organization carrying with them information that only exists in their mind;
understanding how much effort should be placed into producing the data (e.g., for a report on the average number of daily users of an app, a cheap estimate value might suffice, whereas for the cost of parts in manufacturing a precise figure may be needed to ensure profitability), avoiding failures, and guaranteeing service levels.
A small business doesn’t need to invest in expensive software to attain those benefits. A tiny non-profit was able to quickly put together a Google spreadsheet containing basic metadata for its core data assets, covering:
Who (or which system) is responsible for acquiring and processing the data.
Where the data can be found.
Regular reports where the data is used, where they are stored, and who owns those reports so they can be informed if a change is made that could affect the output.
Known issues with existing data assets (e.g., “The total number of active program participants in Spreadsheet X is inaccurate due to a technical issue with the aggregation; for now, if you need this number, it’s better to manually calculate from the subtotals by program found in SpreadsheetY.”)
An unexpected effect of the initiative to build this simple data inventory (and designate an owner to keep it up-to-date) was avoiding the expense of hiring a new employee when it was time to expand the non-profit operations. The previously overworked staff no longer had to spend significant amounts of time tracking down the data and recalculating performance measures that were now maintained in a shared drive, and the same number of staff members was sufficient to support the new activities.
3) Keep looking for low-hanging fruit opportunities to improve the way strategic data is collected and used to enhance performance.
When you know exactly which data assets are critical for the business and what measures need to be continuously tracked, it’s much easier to justify allocating resources to the effort of improving data acquisition and processing over time.
A small business owner needed workers to report daily when they had finished an off-site job. The legacy process required the workers to text the owner when they were about to leave the site, which was very inefficient and error-prone. Workers had to spend time manually writing their messages, and there was no easy way for the owner to confirm if all workers had reported back, or track follow-up actions when someone had to come back to the same location the next day to finish a job.
With the help of a technically-savvy consultant, a few changes were introduced. An icon was added to the workers’ smartphones that when clicked opened an online form where they could simply check a box to indicate “work completed” or select another relevant option before hitting send. The form was linked to a spreadsheet that stored the data. At the end of the day, a simple script produced a summary report for the owner, highlighting the work completed and any deviations from schedule.
The structured work logs created the opportunity for the business owner to start identifying relevant performance measures such as percent of jobs with unplanned delays, rejects and rework, etc. The new spreadsheet inspired the use of a reporting tool to visualize weekly and monthly trends. Now the data supports all kinds of operational decisions, from where to allocate resources to when to start hiring more workers as the demand for the services grows.
# # #
The proposed actions here may seem like a lot of extra work, but this is one of those cases of pay now or pay a lot more later—in lost productivity, duplicative effort, wasted time, second-guessed decisions, and financial or reputational damage from actions taken using incorrect or incomplete information.
Regardless of size, sector, industry, or budget, any organization now has access to numerous tools and resources (many completely free) that can be leveraged to make it easier for both decision-makers and frontline staff to realize value from data. The biggest barrier to eliminating the problems mentioned here is simply that it hasn’t been done yet.
Begin. Take an incremental approach that can show a good ROI at each stage, and soon you should start seeing the compound effect that treating data as a critical asset can bring to your organization’s “decision yield”.

