Teaching Hiring Teams by Doing: Building AI-Powered Practice Activities
Role: Product Lead, end-to-end, from concept through beta to GA launch
Impact: 100,000+ activities completed in the first year; adopted at scale by some of the world’s largest enterprises
The problem
Anyone involved in hiring – recruiters, hiring managers, interviewers – learns their hardest skills on the job, with real candidates as the practice material. Probing a vague answer, delivering difficult feedback, spotting bias in your own judgement: these are skills you only build by doing, and you can’t get better at interviewing by watching someone talk about interviewing.
The insight was simple and old (people learn by doing), but delivering it at scale had always been impossible. Roleplay needs a partner, and human partners don’t scale. Large language models changed that.
The bet
We started this work in 2024, when AI maturity was still relatively low and the pace of innovation was extraordinary; models that led the market at kickoff were obsolete within months. GPT-3.5-turbo was the best model we could afford to run at scale. It was, frankly, not good enough for the experience we envisioned, and running anything better was prohibitively expensive.
So the core product bet wasn’t about features. It was a bet on the curve: that model capability would rise and inference costs would fall faster than our build timeline. We designed the product architecture for the models we’d have at launch, not the ones we had at kickoff. That bet paid off. By the time we shipped, the quality-to-cost ratio had improved by an order of magnitude, and the experience we’d designed around became viable.

What we built
AI-powered activities let learners practise real hiring scenarios in conversation with an AI, then receive coaching on what they did well and where to improve. Every design decision traced back to a small set of principles:
- Hands-on practice: apply skills in realistic hiring scenarios, not abstract exercises.
- Tailored feedback: personalised insights on each attempt, so learners know exactly how to refine their approach.
- A safe environment: practise critical decisions and difficult conversations in a risk-free space, where mistakes are learning moments rather than a bad candidate experience.
- Skill retention: repetition is built in; learners can rerun scenarios until the skill sticks.
- Scalable by design: personalised, hands-on training delivered to thousands of learners at once, something human roleplay could never do.
- Data security first: enterprise customers trust us with their teams, so no personal information is shared with model providers.
Building quality into a probabilistic product
An AI product’s hardest question is “how do you know it’s good?”, not on average, but for every learner, every time.
Our answer was to put evaluation at the centre of the build. Our instructional design team, the people who actually understood what good coaching feedback looks like, built detailed evals for every single activity. This meant AI quality wasn’t an engineering afterthought; it was owned by the same experts who owned learning quality, with pass/fail criteria a model response had to meet before an activity shipped.
What we deliberately cut
Real-time voice was the obvious “wow” feature. Practising an interview by speaking is clearly the end state. We cut it. The latency and cost simply weren’t there yet, and a laggy voice experience would have undermined trust in the whole product. Shipping a text-only experience that worked every time beat shipping an impressive demo that worked sometimes. Voice remains the natural next chapter, on the same cost curve we bet on originally.
Running the beta
For a product this new, the beta wasn’t a QA exercise; it was the product discovery engine, and I ran it as one.
Pitching it. I partnered with our Account Managers and CSMs to recruit 16 companies into the beta, positioning it as a partnership rather than a favour: early access to a genuinely new way of learning, and a direct hand in shaping it. Honesty was part of the pitch. This was early-stage AI, it wouldn’t be perfect, and that was precisely why their feedback mattered. Setting expectations that way turned rough edges into collaboration points instead of complaints.
Consolidating the feedback. With 170+ learners across 16 companies, feedback arrived from every direction: in-product responses, learner comments, and admin conversations. I consolidated it all into a single view, separating what learners said from what the usage data showed, and distilled it into a prioritised set of themes to drive improvements ahead of launch.
The clearest example was activity length. Many activities were taking 15 minutes to complete, and learners told us plainly that this was too long. Our instinct was to aim for 5 minutes, but when we tried, we couldn’t deliver meaningful practice and feedback in that window. We landed on 10 minutes per activity: long enough to add real value, short enough to respect a learner’s attention. It’s the kind of trade-off no amount of internal debate resolves; only real learners could tell us where the line was.
Playing it back. I built a results deck for each beta company and presented it in a wrap-up call: what their learners said, what we changed as a result, and what was coming at general access. Closing the loop did two things. It proved their input had shaped the product, and it converted beta participants into launch champions. Our strongest early adopters at GA were the companies who had watched their own feedback become the product.

The results
Following the beta, we launched to general access in early 2025. From there, everyone in the hiring process – recruiters, hiring managers, and interviewers – completed 100,000+ activities within the first year. The product was rolled out at scale by household-name enterprises spanning global consulting, enterprise software, telecommunications, and industrial engineering – organisations that collectively hire tens of thousands of people a year and hold their training to exacting standards.
The qualitative feedback told us we’d closed the theory-to-practice gap we set out to close:
“Being able to actually practice and receive feedback from the lesson has helped in retaining the information.” – Learner
“It is a great combination of theory and practicing. The AI parts work very good. Would love to see more of that format.” – Company Admin
“This was one of the best SocialTalent trainings I’ve done. The content was great, but the ability to apply it in a relatable scenario made it even better.” – Learner
What I learned
In AI product management, timing your bet on the capability curve matters as much as the product itself. If we’d waited for the models to be ready, we’d have shipped a year behind the market. If we’d shipped what GPT-3.5 could do on day one, we’d have burned learner trust. Designing for the model you’ll have, while ruthlessly cutting what the current one can’t deliver, was the discipline that made this work.
Read more about this on SocialTalent’s website.

