How I Run a Customer Beta: Validating AI Activities with 16 Companies
Role: Product Lead, designed and ran the beta end to end
The subject: SocialTalent’s AI-powered practice activities, tested with 172 learners across 16 customer companies in June and July 2024
Why the beta mattered more than usual
For a conventional feature, a beta answers “does it work?” For an AI product in 2024, it had to answer something harder: “is this actually valuable, and do people trust it?” We were putting generative AI in front of enterprise learners for the first time, and no amount of internal testing could tell us how real recruiters and hiring managers would respond. So I designed the beta as a genuine research instrument, not a soft launch.
Designing the test
Rather than turning the feature on and watching, we built a dedicated test environment inside a real learning path: How to Use Behavioral Interviewing Techniques, instrumented with three short mandatory activities, one longer optional activity, and a reflection exercise at the end.
Two design choices mattered most:
We measured learning outcomes, not satisfaction. The reflection exercise asked learners whether they had improved at the four specific outcomes the activities targeted, such as asking effective behavioural questions and using probing questions to get deeper answers. “Did you enjoy it?” is a vanity question; “can you do the thing better?” is the product’s actual job.
We opened multiple feedback channels. Learners could respond in the reflection exercise inside the path, through a survey shared by email, or directly through their admins. Different people give feedback in different places, and a single channel silently filters out everyone it doesn’t suit.
Recruiting the customers
I partnered with our Account Managers and CSMs to bring 16 companies into the beta, pitching it as a partnership rather than a favour: early access to a new way of learning, and a direct hand in shaping it. Honesty was part of the pitch. This was early-stage AI and it wouldn’t be perfect, which was precisely why their feedback mattered. Setting expectations that way turned rough edges into collaboration points instead of complaints.
Synthesising what came back
Over two months we gathered quantitative outcome data plus 31 pieces of qualitative feedback across all channels. I consolidated everything into a single affinity map, tagging each piece of feedback by theme: content quality, value from practice, AI behaviour, platform experience, accessibility, diversity of scenarios. That tagging turned 31 anecdotes into a prioritised improvement list, and kept the negative feedback visible instead of buried, the red notes sat on the same board as the green ones.
The synthesis also separated what learners said from what the data showed. The clearest example was activity length: learners told us plainly that 15-minute activities were too long. Our instinct was to aim for 5 minutes, but we couldn’t deliver meaningful practice and feedback in that window, so we landed on 10, long enough to add real value, short enough to respect attention. Only real learners could tell us where that line was.

Closing the loop
Every piece of critical feedback got a response, sometimes individually: one learner who felt they hadn’t improved simply wanted more time in the activity, so we extended it, and said so. At the customer level, I built a results deck for each company and presented it in a wrap-up call: what their learners said, what changed as a result, and what was coming at launch. That playback converted beta participants into launch champions; our strongest early adopters at general access were the companies who had watched their own feedback become the product.
What the beta told us
The outcome data was overwhelmingly positive: of the learners who completed the reflection exercise, all but two reported improvement across the four learning outcomes, and both exceptions had clear explanations we could act on. The qualitative feedback confirmed the theory-to-practice gap was closing:
“Being able to actually practice and receive feedback from the lesson has helped in retaining the information.” – Learner
“This was one of the best SocialTalent trainings I’ve done. The content was great, but the ability to apply it in a relatable scenario made it even better.” – Learner
The product launched to general access in early 2025 and passed 100,000 completed activities in its first year.
The compounding value of testing with customers
The beta’s product findings were valuable, but the relationship dividends may have mattered more, and they’re the reason I’d run a customer beta even for a feature I was certain about.
It builds trust before the product has earned it. Inviting customers behind the curtain, showing them something unfinished and asking for their honest reaction, signals confidence and respect. Customers who are treated as partners forgive rough edges that would frustrate them as buyers.
It creates genuine buy-in. People champion what they helped build. By launch, 16 companies didn’t just know the product existed; they had fingerprints on it, and their admins could tell their own learners “we shaped this.”
It strengthens the account relationship beyond the product. Recruiting through Account Managers and CSMs meant the beta doubled as an account touchpoint: our commercial teams got a reason to have forward-looking conversations with customers about where the platform was going, not just how the renewal was tracking.
It de-risks the launch itself. By general access we already had reference customers, real usage stories, and validated messaging, sales enablement built on evidence instead of promises.
None of this happens with an internal QA cycle or an anonymous test panel. The willingness to spend real time with real customers, in their context, on their problems, is what converts a validation exercise into a durable relationship asset.
What I learned
A well-instrumented beta doesn’t need big numbers. Twenty-two reflection responses and 31 qualitative comments were enough to make every major pre-launch decision, because the test was designed to answer specific questions: did learners improve, where did the experience break, and what would make customers champion it. Volume is a luxury; instrumentation is a requirement.

