Skip to content
← All essays

Growth

Growth, Personalization, AI, and Data (A Primer)

AI isn’t making the problems within corporate data systems impossible to ignore yet. But it will soon. As AI fluency spreads and AI-native systems become commonplace, the primary differentiator will be data, both internally/operationally (including in the AI-native system itself) and in the product.

Yes, data is, has been, and always will be imperfect.

No, the currently acceptable level of imperfection will not continue to see success in today’s market.

Let’s break this down.

In this article, I outline common data trade-offs made when scaling AI and the risks and opportunities associated with them. These trade-offs fall into 6 categories: Completeness, Accuracy, Relevance, Connectivity, Legibility, and Privacy.

Introduction

First, let’s establish three key, related assumptions:

  1. Broadly, personalization is a core lever for growth and engagement.
  2. Users’ baseline expectations for products require out-of-the-box personalization; and
  3. This could manifest as any number of solutions, but the most valuable ones are AI-driven (given broad strokes efficiency gains from AI).

There’s more to say on these assumptions. I’ll save that detail for another post.

Bearing these assumptions in mind, business and innovation advances tend to prioritize speed over fidelity. With AI, the trade-offs, risks, and opportunities with this approach usually come back to data (and, in a way, back to my thoughts on context). Personalization-driven growth and engagement experiences will only be as good as their data practice. Many will have an overabundance of opportunity.

Note: I won’t explicitly discuss how to use AI-native systems to address these risks and opportunities operationally, but it should be assumed that such approaches are available to some degree, making these activities more efficient to pursue where they previously were not.

Completeness

Do all records have the same data points (as appropriate)?

  1. In a basic case, does every user record have a first name and a last name?
  2. In a more advanced case, does every user record have associated topics of interest?
  3. In an even more advanced case, does every user who subscribes to our pregnancy journey newsletter have at least one associated child record?*

Customer data platforms (CDPs) are constantly updated with new user data points that can be used as facets in personalization models. If a feature development results in new data capture, it’s common to prioritize capturing data from users who create accounts after the feature is released. Data capture might be included later in onboarding or implemented to be collected in the background only for new users (sometimes intentionally, sometimes as an oversight).

But what about returning users or less engaged users that you’re trying to win back? If your build-out doesn’t include these users, you’re leaving money on the table. Automated backfills and backfill/backfill verification campaigns are known solutions to this issue. Size these sub-opportunities against the same time horizon and see if you still think it’s best to leave them out.

Overall, this risk/opportunity isn’t new/unique to AI, but AI will create the expectation that this gap be addressed (read: closed) sooner than has been the norm.

Note: This is a real example from my time at People, launching a pregnancy journey newsletter program on parents.com (opens in new tab).

Accuracy

Is this data accurate? If you start digging, you’ll find it’s often unclear. That could be because no one knows how the data is processed and/or defined (more thoughts on this below in Connectivity), because it’s not audited regularly, or because it’s not fresh. In turn, the solutions to this are documentation, regular auditing, and features/campaigns that prompt users to confirm accuracy and keep data updated.

A lack of accuracy is particularly damaging when features are built to deploy with high precision (see below on Relevance). Have a personalized campaign for Bushwick residents above 35? Birthdays tend to stay accurate, but too bad half your qualifying users moved to Bed Stuy and just didn’t report it to you. User data collection and management must be developed as twin strategies that work together. If there’s no plan for data refreshment, there’s a plan for silent accuracy drift that will show up in behavioral and revenue metrics, but not in model evals.

If there’s no plan for data refreshment, there’s a plan for silent accuracy drift that will show up in behavioral and revenue metrics, but not in model evals.

A note here on data processing: when organizations look to cut costs, they might perform mass deletions of entire records or specific data points across records. The heuristic often used for this is: Is this data being used now? We need to start asking a different question: Will anyone’s line of business require this data in the next x years? Determining a directional half-life for the data* is a good proxy for x; <x is a good proxy threshold for considering a validation/refresh campaign. Senior ICs and above should all have an opinion on this, per their domain/scope. Think of this as an informed first “gate” on the decision. If there’s even a single yes, compare the cost of running a validation/refresh campaign (before executing the mass deletion) against the cost of completely re-acquiring that data, bearing in mind the half-life calculation. Usually, the former wins out.

A similar approach can be taken when considering deletion of entire records. This is usually a question of whether a cohort of users is planned to be reactivated and whether a user in that cohort has a complete enough record to be reactivated, followed by cost comparison analysis if necessary. An entire user record’s half-life can serve as the proxy threshold for considering a reactivation campaign.

Note: This is a great opportunity to partner with data science on quantitative research. If you don’t have the time or resources to do this for individual data points, a great starting point is to align stakeholders on straw-man half-lives for groups of data points. I say groups because different data behave differently, and even creating a basic heuristic around this goes a long way. For example, data such as name, email address, phone number, and birthday typically don’t change. Their evergreen nature makes them among the most valuable to keep and most costly to replace.

Relevance

Do we have the data we need to achieve the level of precision we’re hoping for in our personalization?

Hint: the more precise you want your personalization to be, the more precise your data must be. You can’t create personalized experiences for ophthalmologists effectively if user records don’t have a job title facet.

This falls into the same data processing trap I covered above. It might sound obvious, but more precise data is more valuable because users generally engage more deeply with more precise personalization. This is an assumption that won’t hold true in every case, but it’s the one I see teams miss more frequently. Compare the savings of a mass deletion against the cost of re-acquisition to determine your approach on this one.

Connectivity

Is the data in all the right places for the user experience to be delivered? Data processing, the data lifecycle, and data transmission to operational systems form a complex web. This is where data risks begin to become more advanced. You’re capturing data and processing it into storage, but you’re not transmitting it to reporting tooling efficiently. Usually that looks like:

  1. you’re not transmitting enough data, or
  2. you’re running into latency issues, or
  3. data’s under-transformed, or
  4. data’s over-transformed.

In my humble opinion, this is an interesting technical area for growth experimentation. Similar challenges are seen with transmission back into production and into operational tooling (I’ve even seen operational tooling only have access to the DOM); this is the more common opportunity space for growth teams. Further, similar challenges are seen with observability for more technical products.

Sophisticated data programs recognize that these challenges are intertwined, not isolated; they all affect personalization experiences and amplify the risks and opportunities in the other categories this article covers.

Stakeholder alignment is the key to moving through this opportunity space. ETL, reverse ETL, storage, and deletion rules need to be determined, documented, and regularly audited cohesively. This includes data summarization and the timeliness of its availability in downstream systems. Aligning on these rules is even more critical when dealing with high volumes of data. It’s another one that sounds obvious, but I’ve seen teams commonly skip it or do it on the fly for certain one-off data.

Legibility

Can a non-technical expert leverage/interact with data in their workflows with confidence in data meaning? In other words, is xdefined the same way across the system, and are x, y, and z available consistently across the system? Even further, is consistently transformed data available across the system (to account for the disparity between transformation functionality within specific tools)?* On the simpler end, is data actually labelled across the system? If you’re primarily answering “no” to these questions, you’re not in a position to effectively evaluate your growth, let alone the effects of personalization on that growth. Extrapolating further, the major opportunity being missed here is attribution of growth to AI-driven product development.†

The most often ignored solutions to this are data labelling and standardized naming conventions (enforced across the system, including in the codebase, with agreed-upon deviations on an as-needed basis). Database management, operational tooling management, and technical integration management are additional key tactics to deploy for more effective success. The data labelling and standard naming conventions are artifact outputs of the same stakeholder management from Connectivity (above). Without this consistency in data legibility, AI-driven features become modern GIGO machines (opens in new tab).

Note: “Consistently,” here means relative to each tool, not the same for each tool. One tool may need partial transformation while another needs full transformation, in order for non-technical staff to effectively use it while handling data. In this same vein, another tool may have full transformation functionality, which allows it to handle data closer to a raw state. And not for nothing, going back to Connectivity (above), there are financial implications of all of this decisioning.

Note: “AI-driven” might mean in the product or in the actual product operations. A deeper dive on attribution as regards CapEx is for another time, but I want to call out that attribution of growth to AI is more complex than it is simple.

Privacy

Data is a team sport, and that’s most apparent in governance, risk, and compliance (GRC). However, it’s not apparent here because people are collaborative, but rather because they often share a frustration with friction in this area and resist meaningful collaboration. Most teams I’ve met would rather take on the risk of being fined (US law, at least, is set up to encourage this; yet another post for another day) than build GRC in from the beginning. This is generally something I try to advise against because the real risk is reputational. If a preventable GRC issue occurs, consumer trust will drop, and they’ll be less likely to provide you with the data you need for effective personalization. Of course, this will negatively impact ROI.

GRC operations are their own separate beast that I won’t tackle here, but I will say that I’ve seen countless poor decisions come from poor translation between product and GRC (here, privacy, security, and legal) teams. Product teams don’t understand how to properly interpret policies and regulatory guidance to assess risk in product workflows and vice versa. The primary role of translation will continue to fall to PMs. Many will be tempted to default to advisement from PMs leading privacy product teams, but I’ve seen many of those such teams where the PM doesn’t inherently need or regularly use GRC knowledge to execute their roles. It’s incumbent upon PMs to develop their own foundational understanding of GRC (and then push deeper) to drive better outcomes, particularly when it comes to AI-driven solutions where regulation feels like the wild, wild west. This aligns with my thoughts on PMs becoming technical.

Conclusion

Completeness, Accuracy, Relevance, Connectivity, Legibility, and Privacy work together to shape data practice within an organization and inform the success (and failure) of AI-driven personalization, which is tantamount to growth. While data is a team sport, I encourage PMs to explore this as a typically underdeveloped area of ownership that they can leverage to drive outsized outcomes.

Again, data will never be perfect, but regularly considering these categories will help teams better estimate risk, size opportunities, and steer outcomes. It’s also worth noting that this applies to most product categories and not exclusively consumer products. Personalization expectations carry over to B2B/SaaS products, perhaps with a slightly different shape. The starting assumptions apply just the same, but the underlying strategy to approach these opportunities will look different.