Steam Market Research 7 archive
Frozen research artifacts, September 7, 2026. Large Parquet datasets, raw API pages and the full announcement inventory remain in the local research directory.
# Language attention on Steam — September 7 UTC / September 6 Brazil Authorized continuation of Steam research direction 6: identify games with substantial review attention outside English and little English-marked response. No causal claim about localization, buyer nationality, game quality or influence. **Completed result:** 456 games in the balanced primary sample, two supplementary observations, and scored-language subsets already available for 200 primary games. One page request timed out after 458 games; collection stopped without retry. The final primary sample uses the first 38 random selections from each of the twelve original strata. No failed game was replaced. All sampling weights were updated to population size / 38. The planned fresh 32-game API profile pass did not run; `available_profiles.py` uses only summaries already returned. See `collection_amendment.json`. All network collection is inactive. ## Sampling The frame is the independently collected September 5 catalog: paid, non-explicit games dated 2019 through August 2026 with at least 100 filtered reviews. There are 13,536 eligible games. Fully delisted/unavailable apps may be absent. Current free status, descriptors and release dates are imperfect historical measures. The frozen sample has 480 games: 40 randomly selected within each of twelve release-period x review-count strata. Periods are 2019–2022, 2023–2025 and January–August 2026. Review bands are 100–499, 500–1,999, 2,000–9,999 and 10,000+. Each member has weight `population_n / sample_n`. Names, descriptions, language support and any observed language shares did not enter selection. The first 24 games are a balanced feasibility pilot; they remain in the frozen sample. The public storefront's review-widget JSON includes global and English-marked Steam-purchase summaries for many larger games. Four independent API checks agreed within a predeclared one-percent or three-review tolerance. Small differences can reflect collection/cache or filtering differences and are preserved. Smaller games generally require API summaries for all and English; larger games with unavailable review props receive the same fallback. Store pages are requested with English language and US country. Individual review rows and account IDs are discarded; no media binaries are downloaded. The sample frame uses September 5 counts; new summaries are collected September 7 UTC. Keep both, with timestamps and count differences. These are a collection window, not an atomic snapshot. Missing values remain missing. A language summary with zero reviews is a measured zero, not a missing observation. ## Questions and thresholds - Fraction of games with at most 10% or 25% English-marked reviews, or a non-English majority. The unit is the game, with sampling weights. - English support versus no listed English support, keeping the latter's small sample explicit. Support flags say nothing about translation quality or when localization was added. - Differences by release period, review band, current price, and overlapping tags. Small tag/domain samples are exploratory, with sample sizes retained. - Discovery screen: at least 500 non-English reviews, fewer than 100 English reviews, and at most 10% English share. This separates relative language concentration from globally well-known games with large English totals. - Within-game English versus non-English recommendation proportions when both sides contain at least 50 reviews. Reviewers in different languages are selected populations; a difference is not a cultural or localization effect. The primary proportions use stratified weighted estimates. Approximate 95% sampling intervals use a design-based linearized ratio variance with finite population correction within the original twelve strata. These intervals do not incorporate source errors, language mislabeling, coverage bias, or temporal drift. Weighted domains use all sampled observations for variance estimation, including zeros outside the domain. Do not present raw sample proportions as population estimates. Review-weighted English share is a separate, secondary statistic and can be dominated by a few very large games. Boundary estimates with no observed variation have no reported interval; the plug-in zero variance is not interpreted as certainty. "English" means the language label attached to a review. It does not identify the reviewer's nationality, actual text language, preferred game language, residence, cultural identity, or first language. Review shares are not sales, players, revenue, awareness outside Steam, or localization return. ## Language profiles A planned second diagnostic sample would select sixteen games from the <=10% English-share bin and eight each from 10–50% and >50%, or all if fewer. Fixed random seeds are saved. These profiles deliberately overrepresent low-English cases and must not be pooled as estimates of language-market size. Each profile obtains all-language and English summaries plus Simplified and Traditional Chinese, Russian, Japanese, Korean, Brazilian Portuguese, European and Latin American Spanish, German and French. The unqueried languages remain an explicit remainder; measured shares are never renormalized to 100%. Language-specific counts are acquired serially and may not reconcile exactly if the source changes or caches differ. All discrepancies are retained. The storefront also exposes some language-specific score summaries. These are a selected subset and never a complete language distribution. They can prove that a named language supplies a majority when its count exceeds half the global count; missing entries cannot establish zero or weak attention. A planned separate adaptive discovery check would take the twenty largest cases meeting the strict discovery threshold, plus up to eight additional low-English-share build-tag examples with 1,000+ total reviews and at most 500 English reviews. Its purpose is to identify a majority review language for the illustrated games. It queries languages in a fixed order and stops once one measured language alone exceeds half the all-language count. Missing languages remain unmeasured. These cases are explicitly selected examples, not an additional probability sample. That script was not run after the timeout; no target manifest or new requests were produced for it. The completed discovery annotations use existing data only. The original sample contains seventeen games without listed English support; none has a wholly missing language-support array. A pre-outcome weighting check against the full frame's known English-support flag is retained in `sampling_benchmark.json`. No sample was replaced because of that comparison. Primary documentation: - https://partner.steamgames.com/doc/store/getreviews - https://partner.steamgames.com/doc/store/localization/languages ## Collection boundaries and reproduction Scoped cap: 1,400 requests including four initial probes, minimum two seconds between request starts, no retries, stop on HTTP or network errors, response size cap 4 MB. The research-root paused bulk-catalog policy remains unchanged. All new calls use a single locked collector. A `STOP` file or inactive scoped policy prevents further requests. Policies are closed at completion. Scripts: 1. `prepare.py` freezes the frame and sample; do not rerun to replace outcomes. 2. `collect_store.py 24`, `pilot_check.py`, then `collect_store.py 480` collect the primary summaries. These require the scoped collection policy to be active and are not needed to reproduce saved results. 3. `analyze.py` joins observations and computes primary weighted estimates. 4. `available_profiles.py` reconstructs the already returned scored-language subsets and annotates discoveries without any network work. 5. `charts.py`, `case_chart.py`, and `validate.py` reproduce the inspected figures and verification. The planned `prepare_profiles.py`, `collect_profiles.py`, `analyze_profiles.py`, and `discovery_languages.py` remain as unexecuted phase provenance and are not needed to reproduce the completed study. Use `../.venv/bin/python`. Requests, sanitized aggregate response records, sample probabilities and specifications remain available. Original raw catalog data and earlier studies are not changed.
# Steam beyond English-marked review attention The study finds a substantial population of games whose Steam review response is weakly represented in English, including games that already list English support. The useful result is a measured discovery set and a map of review attention, not an inference about nationality, sales, or creative ancestry. ## Scope and collection outcome The frame contains 13,536 paid games outside Steam explicit-content descriptors 3/4, released from 2019 through August 2026 and with at least 100 filtered reviews in the September 5 snapshot. This is a metadata exclusion, not an exhaustive manual content rating. The originally frozen sample selected 40 games from each of twelve period/review-count strata, without using names, English descriptions, language support or known language shares. A page request timed out after 458 games, and the collector stopped without retry, at 768 total requests including four initial probes. There was no HTTP 429. The completed primary study uses the fully collected balanced prefix: 38 random games per stratum, 456 total. Two further completed records remain supplementary. The failed game and uncollected games were not replaced, and the primary weights are population stratum size / 38. The proposed fresh 32-game language profile collection did not run after the stop. Nevertheless, 200 primary games already returned scored-language subsets in their storefront review metadata. These support 1,221 English-versus-named-language comparisons with at least 50 reviews on both sides. Their coverage is selected toward larger games and larger review-language groups. Unreported languages remain an explicit remainder, not zero; no complete language census is claimed. The primary sample contains 6,681,746 current Steam-purchase reviews, including 3,417,300 marked English. English and all-language counts come from the same store response or matching API filters. Four pilot store/API checks agreed within one percent or three reviews; small cache/collection differences remain in the raw evidence. Counts were collected September 7 UTC, not atomically. ## How common is low English representation? | Measurement | Weighted game share | Approx. 95% sampling interval | |---|---:|---:| | More than half of reviews are non-English | 49.9% | 43.7–56.2% | | At most 25% English | 21.9% | 16.7–27.1% | | At most 10% English | 13.3% | 8.9–17.6% | | At least 500 non-English reviews, fewer than 100 English, and at most 10% English | 4.7% | 2.7–6.7% | These describe the 100-review-plus sampling frame, not all Steam games. Intervals reflect stratified sampling with finite population correction, not source coverage, review-language labeling errors or changing counts. The last screen corresponds to roughly 634 games in the frame, with an approximate sampling range of 367–901. It separates low relative English share from having an absolutely small English review count. A game can have a low English percentage and still have thousands of English reviews. The secondary review-weighted English share is 50.3% (48.0–52.6% sampling interval). This weights large games much more strongly than the game-level measures. It is not a share of customers, units or revenue. Of 51 observed primary games at <=10% English share, 38 list English support. Weighting gives about 63% of this low-English group with English support. Among supported games, the estimated low-English-share rate is 8.8% (5.2–12.3%). Lack of an English version cannot explain every case. This does not demonstrate ineffective localization: the data do not establish when support was added, translation quality, or how the same game would have performed without it. All fifteen primary games without listed English support have at most 25% English reviews, and thirteen have at most 10%. This is a small subgroup; boundary estimates with no observed variation do not get a spurious 100–100% confidence interval. The full frame independently lists English support for 95.1% of games, so English availability is already widespread in this frame. ## Where the pattern is stronger Using overlapping top-twenty tags, the weighted <=10%-English estimates are: | Tag | Observed games | Low-English cases | Estimated share | |---|---:|---:|---:| | RPG | 138 | 31 | 28.1% (17.5–38.8%) | | Visual Novel | 56 | 21 | 33.0% (18.5–47.5%) | | Strategy | 117 | 19 | 16.5% (7.2–25.8%) | | Simulation | 152 | 22 | 16.2% (8.5–23.9%) | | Adventure | 228 | 29 | 12.9% (7.0–18.9%) | | Action | 229 | 13 | 6.5% (2.1–11.0%) | The RPG and visual-novel contrast with broad Action is worth investigating. These are not exclusive genres or adjusted causal effects. FMV has eight cases among only fourteen sampled games; its high estimate is not a precise population ranking. Deckbuilder and autobattler samples are too small for confident niche comparisons. The complete domain table retains all sample sizes. Some of the RPG concentration comes from visual-novel overlap. Removing Visual Novel/FMV tags from the RPG domain lowers its estimate to 18.8%, with a wide 8.2–29.5% interval. Across the entire frame, excluding Visual Novel/FMV still leaves 9.0% at <=10% English share (5.0–13.1%). Excluding games tagged Sexual Content, Nudity or Hentai leaves 11.3% (6.9–15.7%). The discovery phenomenon therefore extends beyond the narrative and sexual-content-heavy portions of the sample, while the precise genre rankings remain uncertain. Extreme low-English share is less common in the 10,000-plus review band: 2.1%, versus 12.9%, 16.9% and 12.7% in the 100–499, 500–1,999 and 2,000–9,999 bands. Legend of Mortal is a substantial exception. The useful search extends well beyond huge international hits into moderately reviewed games. There is no clear release-period trend in this sample. The <=10%-English estimates are 13.6%, 13.2% and 12.3% for 2019–2022, 2023–2025 and 2026 Jan–Aug, with wide overlapping intervals. These are current language distributions among already-qualifying games, not historical audience shares at launch. ## Concrete discoveries relevant to build-heavy interests These are selected inspection leads, not a representative list of good games. Descriptions below summarize advertised mechanics, not independently played or verified gameplay. Counts are the newly collected primary observations. | Game | Total reviews | English | English share | English support listed | |---|---:|---:|---:|---| | Traveler of Wuxia | 3,148 | 84 | 2.7% | Yes | | Dream of Corpse Lady | 2,080 | 100 | 4.8% | Yes | | Pass the Fear | 2,937 | 178 | 6.1% | Yes | | Fickle Card Legend | 965 | 5 | 0.5% | Yes | | Tetra Project - 原石计划 | 530 | 1 | 0.2% | No | | 轮回修仙路 | 2,943 | 31 | 1.1% | No | | Legend of Mortal | 33,661 | 497 | 1.5% | No | | Eastern Exorcist | 6,099 | 333 | 5.5% | Yes | - **Traveler of Wuxia:** advertises martial-arts card combinations, teammates, events and progression across deaths. Simplified Chinese accounts for 86.7% of reviews, Traditional Chinese another 10.2%. English positivity is 86.9% versus 84.9% in Simplified Chinese: low English attention coexists with broadly favorable English reception. - **Dream of Corpse Lady:** deckbuilding around a puppet army, follower talents, artifacts and power/curse tradeoffs. Simplified Chinese supplies 88.2% of reviews. Positivity is 93.1% there and 92.0% among its 100 English reviews. - **Pass the Fear:** shooting and multiplayer with weapon parts, relic fusion, Tarot cards and interacting build effects. Simplified Chinese supplies 86.1%. English positivity is 87.6%, versus 77.6% in Simplified Chinese. - **Fickle Card Legend:** advertises real-time side-scrolling card battles, Warrior/Mage/Taoist progression and equipment. Five English reviews make it conspicuous, but its overall positivity is only 67.7%. Its dominant language is not established by the summaries collected before the stop. - **Tetra Project:** its Chinese description advertises grid tactics, cards, randomized runs, teammates, interactive environmental objects and extensive mod support. Only one English review was observed. The claimed thousands of mod cards are not a measured count or a claim about viable builds. - **轮回修仙路:** advertises a 3D cultivation roguelike with alchemy, forging, spirit creatures and reincarnation. Activities consume lifespan; subsequent lives retain parts of progression, relationships and equipment. Simplified Chinese supplies 96.4% of reviews. These are developer descriptions, not a reconstruction of the design's origin or proof of its depth. - **Legend of Mortal:** advertises an ordinary sect member whose personality, work, martial arts and decisions affect events and story. It is substantial despite only 497 English reviews and no listed English support. Its language distribution is itself instructive, below. - **Eastern Exorcist:** another larger low-English-share example with English support; 88.3% of reviews are Simplified Chinese. It belongs in the broader action/RPG discovery set rather than only card or simulation categories. Other useful examples broaden the picture: - **Millennium Dream:** 2,588 reviews, 52 English; 95.6% Simplified Chinese. It advertises walking and photography through Chinese Dreamcore/childhood environments. English support is listed, and English positivity is 94.2%. - **Chushpan Simulator:** 2,460 reviews, 75 English; 95.2% Russian. Its pitch describes moving through street life, work and choices in Uryupinsk. English support is listed. - **Love Delivery:** 2,407 reviews, 36 English; 96.9% Korean. A romance/life- management example with English support listed. Its presence demonstrates that the observed low-English cases are not exclusively Chinese-language. Among the eighteen <=10%-English games for which returned data establish a single non-English majority, sixteen are Simplified Chinese, one Russian and one Korean. The other thirty-three lack sufficient breakdowns to identify a majority. This selected scored subset cannot estimate the language composition of all English-light games. ## English and non-English are not homogeneous reception groups Among 389 primary games with at least fifty English and fifty non-English reviews, the median absolute positivity difference is 2.5 percentage points. Forty-one have gaps of at least ten points. English positivity is higher in 239 games and lower in 148, with two ties. These are observed game counts; they are not an estimate that English reviewers are generally more generous. Most paired gaps are modest, but some games look very different by language. Monster Hunter Wilds has: | Review language | Reviews | Positive | |---|---:|---:| | English | 84,802 | 70.4% | | Simplified Chinese | 49,452 | 21.3% | | Japanese | 18,785 | 30.6% | | Traditional Chinese | 15,057 | 45.1% | | Korean | 11,068 | 57.6% | The counts establish disagreement but not whether translation, performance, expectations, pricing, events or another factor caused it. We did not collect review texts to answer that question. Legend of Mortal makes a different point: | Review language | Reviews | Share of total | Positive | |---|---:|---:|---:| | Simplified Chinese | 17,122 | 50.9% | 61.6% | | Traditional Chinese | 11,796 | 35.0% | 93.5% | | Korean | 3,705 | 11.0% | 94.8% | | Japanese | 532 | 1.6% | 98.7% | | English | 497 | 1.5% | 92.4% | Collapsing this into English versus everyone else would conceal a large Simplified/Traditional Chinese contrast. Neither script label is a verified country label. Likewise, English review labels do not establish native English speakers or use of an English game translation. The direction can reverse. Poly TD has 62.5% positivity among 56 English reviews versus 83.9% among 522 non-English reviews. Its English denominator is much smaller, and no specific non-English majority was established. Attention and approval remain separate measurements. ## What this establishes and what remains open The study identifies a measurable population with little English-marked review attention, including accessible English-supported games and several concrete build-oriented references. It does not identify all such games, infer sales or national markets, establish why one language group responds differently, or prove that an unfamiliar pitch represents a new design tradition. Detailed language breakdowns below the storefront's scoring coverage remain incomplete because the planned API profile pass was stopped. That limitation does not affect the measured English-versus-all counts for the 456 primary games. The 200 returned language subsets retain an explicit unreported remainder and are used only where their actual counts support a claim. Source responses, probabilities, code, collection amendment and the selected examples are retained. Both charts were visually inspected. Verification passed 2,134 checks, including identities, weights, source recounts, subset sums, request pacing, the recorded timeout, and inactive collection policies. No individual reviewer records, media binaries, or images were collected as part of the study.

selected_language_profiles.png
