Blog · BL-17

How to structure creator data for pre-campaign due diligence: avoiding follower counts as audience proof

Creator data products should avoid generating black-box authenticity scores. This article discusses how public metrics, authorized analytics, content and brand alignment, disclosure requirements, small-scale testing, and post-campaign feedback form verifiable due diligence.

When brands select creators, follower counts, view counts, likes, and comments are the easiest metrics to obtain. However, what brands actually want to know is whether the audience matches target customers, whether the collaboration content reaches real users, whether the creator fits the brand, and what happens after budget is allocated. A long chain of evidence exists between public metrics and business answers.

A micro-influencer due diligence discussion on r/influencermarketing describes personal experiences of high engagement alongside low click-through and conversion rates, while also outlining various empirical judgments. The poster may have promotional motives, and the thresholds and identification methods in the post lack independent verification. It serves as a checklist of issues rather than training data for a definitive creator authenticity judge.

Start due diligence with brand objectives rather than leaderboards

The same creator might suit product awareness yet fail to drive immediate conversion, cover the right interest groups while having an audience region that differs from sales regions, or possess professional credibility without delivering content at the pace the brand requires. Products first require users to define campaign goals, regions, audiences, content formats, budgets, and restrictions before filtering begins.

It is recommended to convert the brief into verifiable criteria: mandatory, flexible, and optional. For instance, matching content language and sales regions can be mandatory, regular publishing of relevant topics over the past 90 days can be an important condition, and reaching a specific follower scale may be merely optional. Consequently, the system avoids forcing all campaigns into a single generic ranking.

Brand safety should likewise rely on more than keyword blocklists. Teams need to verify recent creator content topics, collaboration density, controversial contexts, and platform status while allowing manual review for edge cases. Machines suit discovering materials that require review rather than applying permanent labels to individuals out of context.

Public channel metrics are signals rather than precise audience profiles

The YouTube Channels resource shows that public subscriber counts round down to three significant figures, and channels can also hide subscriber counts. Furthermore, the statistical definition of view counts may update alongside platform rules. These details illustrate that public numbers maintain display and definition boundaries even when sourced from official APIs.

Public data supports initial observation of content frequency, topics, video performance distributions, and visible interactions, yet it cannot directly prove viewer age, region, purchase intent, or authenticity. A sudden spike in a specific video may stem from topic selection, recommendations, external distribution, or other reasons. Classifying it as anomalous based solely on curve shape turns correlation into asserted fact.

Product interfaces should display values, time windows, and definitions simultaneously. For example, the median visible view count across the last 12 public videos is more interpretable than the average view count and less susceptible to distortion from a single viral hit. If platform metric definitions change, cross-time comparisons must note breakpoints rather than assuming consistent series definitions.

Establish three evidentiary tiers: public, authorized, and experimental

The most useful structure for creator data due diligence groups fields by how evidence is acquired rather than blending all attributes into a single composite score.

  • Public evidence: channel profiles, public content, visible interactions, publication timing, and public disclosures. Suitable for discovering candidates and checking content alignment.
  • Authorized evidence: platform analytics, media kits, or historical collaboration reports provided through creator authorization. Suitable for verifying audience regions, traffic sources, and granular performance data.
  • Experimental results: brand-owned links, promo codes, landing pages, sales, or incrementality test data. Suitable for assessing the true effect of a collaboration on the brand.

The YouTube Analytics dimensions documentation lists reporting dimensions such as country, city, traffic source, and demographics, while noting that these metrics belong to authorized channel or content owner analytics capabilities. They cannot be retrieved as arbitrary third-party fields via public scraping. Due diligence products must clearly distinguish between publicly readable data and data requiring creator authorization.

Authorized data is also not absolute truth. Screenshots may lack dates, filter conditions, or full pages, and exported metrics might misalign with campaign goals. A more reliable workflow requires fixed time windows, metric definitions, and source proofs. If only screenshots are available, record their scope and unverified items without processing them into precise forecasts.

Avoid selling an inexplicable authenticity score

Follower growth surges, repetitive comments, and unusual interaction distributions can all serve as review signals, yet each has normal explanations. Synthesizing them into a score of 87 and labeling anyone below 60 as fake simultaneously generates false positives, appeals, and legal risks while leaving users uncertain which part to trust.

A better deliverable is a due diligence card, which can contain:

  1. Brief matches and mismatches, with corresponding evidence for each.
  2. Time windows, distributions, and apparent anomaly candidates for public metrics.
  3. Acquired authorized data, missing key fields, and verification timestamps.
  4. Content and collaboration history segments requiring manual inspection.
  5. Recommended small-scale test designs instead of definitive performance guarantees.

Systems can state that visible interactions over a recent period concentrate in a few pieces of content and recommend verifying authorized reach data, but they cannot conclude that interactions were purchased based solely on that finding. For creators themselves, correction entry points should be provided since account merges, viral content, platform campaigns, and content deletions can all explain metric shifts.

The internal creator data capabilities guide explains that creator metrics must distinguish between public profiles, content performance, and unverified inferences. This article extends that same principle to campaign workflows without claiming that public data from any platform can identify actual buyers.

Incorporate compliance disclosures into pre-campaign checklists

Campaign due diligence evaluates not only whether a collaboration is worthwhile, but also whether it fulfills disclosure requirements. The FTC social media disclosure guide for influencers states that when financial compensation, free products, employment, or other material connections exist between brands and creators, endorsement content must make that relationship clearly visible.

Products can integrate disclosure requirements into briefs, contracts, and pre-publication checklists while saving brand confirmations and creator submission records. They should not treat generic templates as universal legal advice or declare compliance complete merely by scanning for a tag after publication. Different jurisdictions, platforms, and content formats may impose distinct requirements that professional personnel should review when necessary.

Prior disclosures within public content can serve as process maturity signals, but a single omission should not automatically trigger a permanent blocklist. Record specific content, timing, and visible evidence to let brands decide whether to continue discussions.

The most reliable score comes from a brand's own small-scale tests

Due diligence reduces uncertainty rather than prophesying outcomes. For candidates with insufficient evidence yet matching content, teams can run limited-scope, clearly defined small-scale collaborations using brand-controlled links, landing pages, or discount identifiers while pre-defining primary goals across exposure, visits, sign-ups, purchases, or content reuse.

Test results should account for costs and time windows while retaining attribution limits. A lack of clicks does not necessarily mean content failed to build awareness, and generated purchases do not stem entirely from creators. Without control groups or incrementality designs, document results as observable correlations rather than rewriting them as definitive causal lifts.

Historical collaboration data from brands can subsequently feed back into filtering, provided it is grouped by campaign type, region, and product. A creator's performance on low-priced products cannot be automatically extrapolated to high-priced services, nor can another brand's case study replace internal experimentation.

Decision records and collaboration form the core of productization

Due diligence products should allow marketing, legal, finance, and operations teams to collaborate on a single candidate record: who nominated the candidate, which data is authorized, who reviewed content, whether contracts and disclosure requirements are confirmed, who approved the budget, and when post-campaign results are logged. Data collection represents only one phase.

Permissions require granular controls. Creator-authorized audience data, contract pricing, and brand conversion data should not be open to all users by default, and public content used for filtering must retain its source and verification timestamp. After campaigns conclude, teams must specify which raw data is retained, which derived conclusions expire, and how creators can request corrections.

Technically, teams can reference the internal unified data API design guide to partition profiles, content, evidence, campaigns, and decisions into distinct objects, while using the data source selection guide to separate public discovery, authorized data, and actual delivery evidence.

An MVP does not need to cover millions of creators

An initial version can serve a single vertical category and a single real campaign: start from dozens of candidates, recording manual filtering reasons, authorized data acquisition rates, brief mismatch reasons, final selections, pre-publication blockers, and post-campaign results. Every rejection should have an explainable reason rather than leaving behind only a low score.

Acceptance metrics include time required to find suitable candidates, manual review proportions, missing authorized data items, correction frequencies for due diligence conclusions, and post-campaign predictive signals. Avoid replacing customer results with database scale, collected field counts, or fraud-detection figures.

This article discusses creator data productization methods and does not imply that EveryInfra has released a creator due diligence service. A quality product avoids pretending that public data can see through an account; instead, it presents evidence of varying strengths to users, enabling brands to progressively reduce uncertainty using authorized data and internal experiments.