8 questions. 3 minutes. Instant reliability score with actionable findings.
No sign-up required
5+
Reliability dimensions
3min
Average completion
0–100
Scored instantly
Part 1 of 2 — Reliability Dimensions
Rate your AI bot
Select the option that best reflects your bot's current behaviour.
How accurate are your AI bot's responses?
ⓘHow often does your bot produce factually correct responses? Include hallucination frequency in your assessment.
Q1 / 8
Likely incorrect
Some errors
Mostly correct
Very accurate
Fully validated
High riskNo risk
Please answer this question to continue.
Does your bot admit when it doesn't know something?
ⓘAgents that always sound confident — even on uncertain or out-of-scope topics — are harder to trust and more likely to mislead users. The best agents say "I'm not sure" or "I don't have enough information" when appropriate.
Q2 / 8
Never
Rarely
Sometimes
Usually
Always
High riskWell-calibrated
Please answer this question to continue.
How completely does your bot address what was asked?
ⓘDoes the bot fully address what was asked? Partial answers can leave users with incomplete or misleading information.
Q3 / 8
Very incomplete
Partial
Adequate
Detailed
Comprehensive
High riskNo risk
Please answer this question to continue.
Does your bot back up its answers with reasoning or references?
ⓘDoes your bot back up its answers with reasoning, references, or citations? Unverifiable responses undermine trust in high-stakes contexts.
Q4 / 8
None
Weak
Basic
Clear reasoning
Strong + references
High riskNo risk
Please answer this question to continue.
Performance — how quickly and consistently does your bot respond?
ⓘMeasures response latency and consistency. Slower agents reduce user experience and limit scalability in production environments.
Q5 / 8
> 10 sec Poor
5–10 sec Slow
3–5 sec Acceptable
1–3 sec Fast
< 1 sec Excellent
High riskNo risk
Please answer this question to continue.
Part 2 of 2 — Risk awareness
These 3 questions don't affect your score — they flag untested risk areas that appear in your report.
Out-of-scope handling tested?
ⓘHave you observed how your bot responds to queries clearly outside its purpose — does it decline gracefully or attempt to answer?
Q6 / 8
Yes — tested
No — not tested
Not sure
Please answer this question to continue.
Security & adversarial testing done?
ⓘHave you tested how your bot handles attempts to manipulate, jailbreak, or extract sensitive data? This is sometimes called red teaming or security testing. Critical before any production deployment.
Q7 / 8
Yes — tested
No — not tested
Not sure
Without security & adversarial testing you have no visibility into whether your bot can be manipulated into producing harmful, biased, or policy-violating responses.
Please answer this question to continue.
Accuracy consistency verified?
ⓘHave you run the same query multiple times to verify your bot's accuracy stays within acceptable bounds across iterations?
Q8 / 8
Yes — verified
No — not tested
Not sure
LLM outputs are naturally variable — without testing you don't know if accuracy stays within acceptable bounds across repeated queries.
Please answer this question to continue.
Your report is ready
Get your full AI Bot Reliability Report
Enter your details to view your scorecard instantly — and receive a PDF copy in your inbox.
Your score preview
—
/100
—
🔒 Enter your email to unlock the full report
✓ Instant on-screen results
✓ Full PDF report to your email
✓ Dimension breakdown + risk flags
✓ Actionable QE recommendations
JG
Jade Global
AI Quality Engineering
Almost there
Where should we send your report?
Please enter a valid email address.
Please enter a valid phone number.
🔒 No spam, ever. Your details are used only to deliver this report and occasional QE insights from Jade Global.
Your reliability scorecard
74
/ 100
Good reliability
Your bot performs well across most dimensions. Minor targeted improvements will get it to production-ready.
Good — minor improvements needed
Sample scorecard
Based on demo answers — take the real assessment for your actual score
Your report
Risk band
0–40
High risk
Likely hallucinations
41–70
Moderate
Needs validation
71–90
Good
Minor improvements
91–100
Strong
Production-ready
Score breakdown
Accuracy—
Confidence—
Completeness—
Sourcing—
Performance—
Key findings
Untested risk areas
✓
Out-of-scope query handling
Good — you have tested out-of-scope handling.
Tested — no flag
!
Security & adversarial testing
Critical gap — no red team testing detected.
Not tested — action required
?
Accuracy consistency across iterations
Untested — accuracy variation across repeated queries unverified.
Not sure — unknown risk
Your personalised action plan
Based on your specific answers — highest priority first
No critical gaps detected across your answers. Focus on maintaining and expanding coverage over time.
Want this done automatically?
Jade Global's AI evaluation platform runs these checks daily against your live bot API, tracks quality over time, and flags regressions before your users notice them.