A 30-Day Games Pilot Gives You 200 Sessions Per Title. Rank the Licensor, Not the Games.
Before you license HTML5 games, run the pilot on delivery, support and licence terms โ not on retention curves your pilot traffic is far too thin to measure.
The request is reasonable and it comes up in almost every catalogue deal: give us fifty titles for thirty days, we will put them live on a test page, and we will decide from the numbers. Nobody argues with it. It sounds like due diligence, it survives the procurement review, and it delays a signature by exactly one month.
The problem is not the pilot. It is what the pilot brief asks the pilot to answer. Most of them are written to rank games, and a thirty-day trial on a test page cannot rank games. It can, if you write it properly, tell you something far more useful โ whether the counterparty on the other end of the licence is one you want to be attached to for three years.
๐งฎ The Number That Decides What a Pilot Can Prove
Start with arithmetic rather than opinion. Comparing two conversion or retention rates at the industry-standard 95% confidence and 80% power needs roughly 16 ร p(1โp) รท dยฒ users per arm, where p is the baseline rate and d is the absolute difference you want to catch. Adobe's own sample-size guidance for Target recommends exactly that 95%/80% pairing, and adds a two-week minimum so day-of-week effects wash out.
Put a real baseline in. GameAnalytics' 2026 Mobile & PC Gaming Benchmarks puts the median mobile Day 1 retention at around 22%, with the median Day 7 just under 4% and Day 30 below 1%. Web behaves differently from mobile, but 22% is a defensible baseline to reason with.
- To detect a 10% relative difference in D1 between two titles โ 22% versus 24.2% โ you need about 5,700 players per title. Roughly 11,400 to separate one pair.
- To detect a 25% relative difference โ 22% versus 27.5% โ you need about 910 per title.
- To detect a gap as wide as 22% versus 32%, about 320 per title is enough.
Now the pilot side of the ledger. Say your test section does 20,000 sessions in thirty days, which is a healthy result for a soft-launched page on an existing site. Spread across fifty titles that is 400 sessions per title; across a hundred, 200. And a session is not a player: Poki's 2026 State of Web Gaming report finds players move through two to three individual titles inside a single session, so those 200 sessions represent fewer distinct people than the number suggests, distributed unevenly because your homepage put six titles above the fold and the rest below it.
So the pilot can catch a title that is catastrophically worse than the others. It cannot tell you whether title 14 beats title 31. Every pilot scorecard that ranks fifty games by D1 and cuts the bottom twenty is ranking noise, then deleting inventory on the strength of it.
๐ Two 2026 Reports, Three Times Apart at Day 7
Even if you had the volume, you would still need a pass mark โ and the published pass marks do not agree.
GameAnalytics' 2026 report puts the mobile median at roughly 22% D1 and just under 4% D7. A 2026 benchmark round-up from Segwise, drawing on Mistplay's genre data, gives puzzle titles 31.85% D1 and 12.18% D7, and hyper-casual 29.31% D1 and 5.90% D7. Day 7 in one source is around 4%. Day 7 for puzzle in the other is over 12%.
Both can be right. One is a median across everything an SDK sees, including the long tail of games nobody plays twice; the other is a genre cohort inside a rewarded-play network with its own acquisition dynamics. They are different populations, not different truths. The GameAnalytics report is explicit that genre-level cuts are unavailable in this edition at all, which should tell you how unstable the categories are underneath.
The consequence for your pilot is blunt. A licensed catalogue returning 9% D7 in a thirty-day test is either a strong result or a failing one depending entirely on which PDF the person writing the summary opened. If the pass mark is arbitrary and the sample is too small to hit it reliably, the pilot is not a measurement. It is a coin toss with a spreadsheet attached.
โฑ๏ธ Thirty Days Is a Season, Not a Sample
Timing distorts what is left. A pilot that runs across a school holiday, a national holiday week, a Ramadan window or a major sports final is measuring the calendar as much as the catalogue. Adobe's two-week floor exists to average out weekday-versus-weekend behaviour on ordinary web traffic; game traffic is more cyclical than that, not less.
There is a second distortion that pilots rarely account for. Poki's 2026 report found 90% of web game players are doing something else at the same time โ music, streams, social โ and only 44% say the game holds their primary attention. Engagement metrics collected under those conditions are noisy by nature, which raises the sample you need rather than lowering it.
And a pilot audience is not your audience. It is whoever found a page you have not promoted, or a slice of your existing traffic redirected into a section that has no navigation, no SEO history, no email flow and no returning-player base. Whatever number that population produces, your live portal will not reproduce it.
โ What a Pilot Actually Proves โ and It Is Worth More Than a Retention Curve
Everything above argues against measuring games. None of it argues against running a pilot. A thirty-day trial is an excellent instrument for testing the supplier, and supplier risk is the risk that actually costs you a quarter. Point it at questions where n=1 is a valid sample, because a single observation genuinely answers them:
- Delivery shape. Do builds arrive as ZIPs, as hosted URLs, as a feed, or as a spreadsheet and a promise? How long between the request and the files?
- Build hygiene. Relative asset paths, no hard-coded absolute domains, no sitelock pointed at the wrong host, no external calls you did not expect. One title tells you the house standard.
- Metadata completeness. Titles, descriptions, categories, thumbnails at the sizes your grid needs, orientation flags, control scheme, language list. Missing metadata is a content project you will pay for later.
- The replacement test. Report one genuinely broken title. Time the response, and see whether you get a fix, a swap, or silence. This is the single most predictive thing in a pilot.
- Compliance artefacts. Age ratings, content descriptors, a data-processing agreement, a statement of what the builds collect. Ask during the pilot, not during the security review.
- Commercial paperwork. An invoice, a reporting sample, a named contact who answers in under a day.
Every one of those is answerable in thirty days with fifty titles and no statistics whatsoever. Every one of them predicts what year two of the relationship feels like.
๐ A Thirty-Day Pilot Plan Built Around Delivery
Structure it in four weeks, with a decision at the end of each.
- Week 1 โ receive and inspect. Take the builds. Do not put anything live. Run every title on a low-end Android handset and a desktop browser, check load weight, check for outbound requests, check the metadata against your schema. Record the count of titles that need any manual fix.
- Week 2 โ integrate honestly. Put them into a real page in your real stack with your real ad or billing layer, not a stripped test harness. The friction you find here is the friction you will pay per title across the whole catalogue.
- Week 3 โ break something on purpose. File the replacement request. Ask for a title in a category the pilot batch missed. Ask a licence question in writing โ territory, sublicensing, hosting rights โ and see how fast a clear answer comes back.
- Week 4 โ read the traffic qualitatively. Look for catastrophic failures only: titles that fail to load, titles with a sub-five-second average session, titles that break on mobile. Do not rank the survivors.
Write the exit criteria before week one, and make them binary. Fixes needed on fewer than X titles. Replacement delivered within Y working days. Licence questions answered in writing. Those are pass/fail. A retention target is not.
๐ข Getting the Performance Answer Without Pretending You Measured It
You still want to know whether the games perform. Three routes get you closer than a thirty-day test page, and none of them requires you to pretend 200 sessions is a sample.
- Ask the licensor for aggregate data across their base. A licensor distributing the same titles to many operators has volume you will never generate alone. Ask what the top decile looks like by category and by region, and ask what the definitions are โ plays, sessions, uniques and starts are four different numbers.
- Decide at category level, not title level. You have enough traffic to tell whether puzzle beats arcade on your audience. You do not have enough to tell whether one match-3 beats another. Buy breadth in the categories that win, and let the individual titles rotate.
- Stage the rollout instead of gating it. Sign for a first tranche, go live properly with promotion and navigation behind it, and take a second tranche on the strength of ninety days of real traffic. That is a measurement. A pilot page is not.
This is also the argument for licensing a catalogue rather than buying titles one at a time. Individual purchases force a per-title judgement you cannot support with data. A catalogue lets you be wrong about any single game without it mattering, and gives you enough inventory that rotation โ not selection โ becomes the lever you pull.
๐ซ Five Ways a Games Pilot Wastes a Quarter
- Ranking fifty titles on thirty days of traffic and cutting the bottom twenty. You have deleted inventory based on variance, and you will never know which good titles went with it.
- Running the pilot on a page nobody can find. No navigation, no promotion, no returning users โ then reading the retention curve as if it described your business.
- Testing in a sandbox that is not your stack. The integration cost you were trying to discover is precisely the cost you engineered out of the test.
- Never filing a support ticket. The pilot ends with a warm feeling and zero evidence about what happens when something breaks at 200 titles instead of fifty.
- Leaving the licence questions until after the pilot passes. Territory, hosting rights, sublicensing and exit terms decide whether the deal works at all. Discovering a blocker in week nine wastes the whole exercise.
๐ฏ What a Licence Covers When You License HTML5 Games Direct
A licence from a direct licensor is a supply arrangement, not a file transfer, and that is what a pilot should be testing. Forestry Games has operated since 2017 and licenses a catalogue of 1,049 titles spanning HTML5 and Android APK builds, with source available where the title and the deal support it. One conversation covers both formats, so a portal and an app do not become two separate procurements.
Practically, that means HTML5 builds you can host yourself or run from hosted URLs, APK builds for Android distribution, branding and white-label options where you need the surface to be yours, and the metadata and assets a real storefront needs. The catalogue's depth is what makes the pilot advice above workable โ you can rotate, re-weight by category and replace weak titles without renegotiating, because breadth is already in the licence. If you are scoping a first tranche, browse the catalogue to see the category spread, or start from the format you actually need and license HTML5 games for web surfaces and buy Android games for APK distribution.
๐งธ Licensing Branded Games for Campaigns, Portals and Events
Some pilots are not really catalogue pilots โ they are a client asking whether you can put a recognisable character in front of an audience. Forestry Games works with branded IP and has brand partnerships including Disney, Nickelodeon, Cartoon Network and Warner Bros, and businesses can license branded game content through it for campaigns, portals, events and apps.
Branded work runs on a different clock from a generic catalogue tranche, because rights and approvals sit upstream of the build. Raise it at the start of a pilot rather than at the end, so the scope conversation and the licence scope conversation happen together. The next step is concrete: ask for a licence scope covering your territory and surfaces, or request a white-label portal demo and put the delivery questions above to it directly. A membership plan is the route if you want ongoing catalogue access rather than a one-off tranche.
๐งญ Rewrite the Pilot Brief Before You Send It
Open the pilot brief sitting in your drafts and look at what it asks for. If the success criteria are retention thresholds and a title ranking, the exercise will produce a number, that number will not mean anything, and the month it costs you buys nothing except the appearance of rigour.
Swap them out. Fixes needed per hundred titles. Days to replace a broken build. Whether the licence questions came back in writing. Whether the integration took hours or weeks in your actual stack. Then take the performance question off the pilot entirely and put it where it belongs โ a staged first tranche, live properly, measured over ninety days with traffic you control. Send that brief, and thirty days will have told you something you can act on.


