Clustering — grouping data points based on characteristics they have in common — can be a valuable tool to uncover hidden associations within data that might not be directly visible to the naked eyes.
Clustering is particularly useful to reveal natural but not necessarily intuitive groups of objects of interest that share a set of common or similar characteristics. In business contexts, we often see this technique used in customer segmentation, when a customer base is divided into groups of people based on factors like demographics, preferences, and buyer behavior.
Used for customer segmentation, clustering can help prevent mistakes such as:
Targeting customers by age, gender, and income in advertising when the most relevant factors defining a company’s “ideal buyer” are education level and whether the person lives in a high-density urban area.
Concluding that an online course with high dropout rates doesn’t work when in reality it provides excellent results for the subset of students with access to a tutor to answer questions and feedback for each lesson received within 2 days of homework submission.
Here’s a good example of clustering used to develop an evidence-based understanding of the customer profiles of a multi-unit retail organization that sells makeup and skincare products:
MECCA created a de-identified dataset of 50,000 customers’ transactions, and fed them into Affinio’s Augmented Analytics platform, which can analyze a dataset based on thousands of attributes simultaneously, identify statistically relevant commonalities, and surface otherwise hidden distinct clusters or communities within an audience.
Within hours, Affinio identified six preliminary audience clusters with distinct and insightful characteristics. (Source)
One can think of many applications for this kind of customer segmentation, including:
Better marketing campaigns, with distinct messaging and frequency of communication for Occasionals vs. Volume Buyers.
Better decisions about product portfolio: ensure that products attractive for Big Spending Regulars are always in-stock.
Customized shopping user experience: product recommendations for Quarterly ‘Replenishers’ may need to focus on good substitutions when their favorite products have been discontinued, whereas for Occasionals and Moderates they may focus more on up-sell and cross-sell.
While in marketing and sales it’s relatively common to see clustering used to improve business results, businesses often fail to exploit its full potential.
How clustering can improve business results outside of customer segmentation
Consider an IoT provider that supplies asset tracking devices to manufacturing and construction sites. Asset trackers are placed in expensive equipment with the dual purpose of preventing theft and time-wasting trips around the site to find equipment that’s moved around.
The device sensors must have their battery replaced from time to time. Replacing the batteries too soon increases labor costs and negatively impacts the environment, while waiting too long causes trackers to stop to report asset location. Clustering can help solve this optimization problem by uncovering the hidden associations between sensor measurements that report on local temperature, signal strength, signal-to-noise ratio, average asset idle time, etc.
Sensor data accumulated over time could be leveraged to produce a cluster like this:
Here, “Regulars” are sensors operating under normal circumstances and experiencing the standard battery lifespan. Other clusters include sensors operating under conditions of low temperature (shorter lifespan), and higher signal-to-noise ratio or idle time (higher lifespan).
Based on this segmentation, the company can make informed decisions that might include:
Advising customers to use a different kind of sensor or plan for more frequent battery replacement when operating in environments with low temperatures.
Reducing the frequency of battery replacement in assets operating under conditions of high idle time or high signal-to-noise ratio to take advantage of the the higher battery lifespan achieved under those circumstances.
There is a vast array of business problems that clustering can help solve
Clustering should be considered any time an organization can benefit from learning which subsets of objects are highly similar to each other, or which data attributes tend to occur together. Other practical examples include:
Classify college students from a university into unique subtypes so that distinct groups with high risk of dropping out can receive customized assistance to increase graduation rates.
Capture the relationships between hundreds of bill of material items spread across multiple locations to improve inventory management in a supply chain network.
Detect misbehaving servers in a network when the use of threshold alerts would set off too many false alarms due to metrics prone to spikes and/or fluctuating baselines.
For certain problem types, clustering may offer a superior solution relative to other machine learning techniques or traditional analytics. Creative uses of clustering may turn out to be the key to protecting your company’s competitive advantage, enhancing its analytics capabilities to derive insights, translate the insights into actions, and drive continuous improvement and sustained impact.



