Published research protocol · Version 1.0

Benchmark methodology

How Synta selected 50 brands, collected 250 shared evidence packs, ran 750 standardized tests, scored visibility and audited the results.

The study asks how visible China-origin B2B brands are to US buyers using AI-assisted vendor discovery, and how those brands compare with established global suppliers and representative challengers.

50
brands
5
industries
250
evidence packs
750
answers

1. Sample design

Five industries contain ten brands each: seven China-origin or China-export brands with English-language and US-market evidence, two established international benchmarks, and one representative international challenger. The sample therefore contains 35 China-origin brands and 15 international benchmarks. It is a purposive commercial sample, not a market-share census.

  • Outsourced biopharma R&D and manufacturing
  • Nutraceutical ingredients and supplement CDMO
  • Industrial and collaborative robotics
  • Stationary battery energy storage systems
  • Cross-border ecommerce logistics and US fulfillment

2. Prompt structure

Each industry-model cell contains 50 independent answers. Two unbranded buyer prompts are each repeated in ten fresh, stateless sessions. Every one of the ten brands then receives three branded prompts: comparison, alternatives, and trust/evidence. This produces 250 answers per model family and 750 answers in total.

Prompt family per industry and modelAnswers
Unbranded category recommendation × 10 independent runs10
Unbranded US procurement scenario × 10 independent runs10
Branded comparison × one per brand10
Alternatives × one per brand10
Trust and evidence × one per brand10

3. Models and retrieval controls

The completed benchmark used Perplexity Sonar, Qwen 3.7 Flash and DeepSeek V4 Flash. Model identifiers, timestamps and output metadata were retained internally. Results describe these model snapshots during the August 14, 2026 execution window.

  • English prompts with a US B2B procurement context.
  • A fresh stateless request for every answer.
  • One shared web evidence pack for each of 250 prompt instances, reused across all three model families to reduce source-set variation.
  • Search location set to the United States with current web evidence enabled.
  • Retries limited to network, rate-limit and provider errors, never because an answer was unfavorable.
  • A hard USD 30 budget ceiling; actual model and retrieval spend was approximately USD 7.15.

4. Scoring framework

The overall score is descriptive and fully decomposable. Component values remain available in the internal research workbook so readers can compare individual signals or apply different weights.

Unbranded recommendation frequency

30 points

Share of unbranded industry runs in which a brand is presented as a viable supplier.

Recommendation prominence

20 points

First-mention position and primary, secondary, caveated or non-recommended placement.

Branded comparison strength

20 points

Credibility, procurement fit, balanced strengths and material limitations.

Citation coverage and source quality

15 points

Claim citation coverage and the mix of official, regulator, research, trade, media, review and community sources.

Claim support and verifiability

10 points

Whether trust-prompt claims are supported by cited or registered official evidence.

Sentiment and procurement confidence

5 points

Positive, neutral, cautionary or negative recommendation posture.

5. Citation and claim audit

Every completed answer contained at least one parsed source. The final dataset contains 3,032 citation instances across 313 unique domains. Domains were classified as official company, government or regulator, academic or research, news or business media, marketplace or review, social or community, trade publication, or other web source.

Trust-prompt claims were compared with the cited source and registered official evidence. Unsupported, contradicted, stale and ambiguous claims were flagged. This is an evidence audit, not a certification of supplier quality.

6. Quality assurance

  • 750 of 750 planned answers completed successfully.
  • All 750 answers had a normal stop condition and at least one parsed citation.
  • No answer remained truncated after repair and re-checking.
  • A stratified semantic audit reviewed 150 answers and 4,500 classification fields.
  • Automated and semantic review agreed on 91.8% of audited fields; disagreements were retained for traceability.

7. Search-to-AI gap

A brand is counted as retrieval-visible when its official domain appears in the shared evidence packs. If that brand receives no unbranded recommendation, the study labels it a retrieval-to-recommendation gap. This is deliberately described as a retrieval-visibility proxy, not a claim about Google or Bing organic ranking.

8. Interpretation and limitations

AI answers vary by model, date, retrieval index, geography, provider policy and prompt wording. A controlled execution window cannot represent every buyer conversation. Brand absence is not evidence of weak products or operations. The benchmark diagnoses machine-visible evidence and recommendation behavior at a documented point in time.

Standardized API outputs are research snapshots and may not be identical to consumer chat-product interfaces. Public attachments contain derived brand metrics and source summaries; full provider answers are retained for internal reproducibility and audit.

Reproduce or challenge the findings

Download the public 50-brand attachment, review the scoring definitions, and contact Synta for a category-specific benchmark or methodological query.

Return to the benchmark report