THE AI AGENT FIELD GUIDESOURCES FIRST. CLEARER CHOICES.
REPRODUCIBLE EVALUATION

Test the work, inspect the evidence.

Inspect AI product test plans, recorded public feature observations, local runtime results and separate source/page checks. Every result states its tested scope and limitations.

What has actually been checked?

60 product plans

Every profile has a specific controlled input, request, expected output and boundary case.

60 profile checks passed

These verify uAgentKit’s rendered content, sources and evaluation sections.

126 official URL checks

HTTP status and readable content establish page access; claim truth and product quality require a separate review.

3 local runtime records

Only the recorded deterministic runtime scope is established. Vendor-account performance tests have not been completed.

3 public feature records

0 passed the selected checks · 1 partial · 2 blocked before output. These are separate records; the two product-account cases per profile remain unexecuted.

Report generated 2026-10-02T05:48:30.286Z. The last recorded run is shown; a source becoming unavailable later does not automatically revoke or validate a vendor claim.

Public feature results and access limits

These records describe actual public-interface inputs and observations within each stated scope. They do not complete the two planned product-account cases or establish overall quality, costs or business outcomes.

Pikpop

Public feature blocked before output · 2026-10-02T05:25:54.349062+00:00

Scope: Official, account-free headphone redesign demonstration using preset demonstration data. Public page navigation and one reset action only.

Conclusion and limits: The execution area remained empty at step 1/5 with zero deliverables after reset. No custom prompt, production model response, supplier quotation or procurement action was produced.

0 passed · 0 failed · 1 blocked checks. Read the actual inputs and observations →

Download the public feature record (JSON) →

Typefully

Public feature blocked before output · 2026-10-02T05:29:00.995097+00:00

Scope: The official account-free AI Text Generator, using its displayed GPT-4o mini and Short settings with one synthetic product-copy request. This does not test Typefully workspace, MCP, scheduling or publication.

Conclusion and limits: The request produced no visible generated text. The browser reported that Turnstile verification timed out; no CAPTCHA was solved. Product factuality and the planned second boundary prompt could not be assessed.

0 passed · 0 failed · 1 blocked checks. Read the actual inputs and observations →

Download the public feature record (JSON) →

Postiz

Public feature partially checked · 2026-10-02T05:33:01.776Z

Scope: Two prompts in the official account-free YouTube Title Generator: an accurate mug-cleaning tutorial and missing profitability-data boundary. This tests the public utility only, not Postiz Smart Agent, calendar, MCP, approval enforcement or publishing.

Conclusion and limits: Both prompts returned titles. The first failed the factual boundary by adding an unsupported 2024 date, a PROVEN claim and a completion-time promise (in Minutes). The second returned five titles with no invented financial figures or profit guarantee, passing that limited financial-claims check, but titles 1 and 4 did not explicitly focus on collecting missing cost/revenue data and failed that instruction. Its SEO Tip is an unmeasured suggestion, not evidence of ranking or views. Overall: one limited condition passed and two failed; neither prompt fully satisfied all requested constraints.

1 passed · 2 failed · 0 blocked checks. Read the actual inputs and observations →

Download the public feature record (JSON) →

Three evidence layers

  1. Product behavior: run the listed cases in the actual product with authorized input and inspect output quality, costs, permissions and failures. Public feature records have the limited scope shown above. Every planned product-account case here remains not executed.
  2. Official documentation: inspect source content and current access separately. An HTTP 200 response does not establish efficacy, current pricing or legal compliance.
  3. This website: check that each profile renders its factual fields, source references, workflows, FAQs and evaluation cases. These tests evaluate uAgentKit.

A local graph/CLI check, where recorded, verifies the named runtime behavior only. It cannot support claims about language-model accuracy, hosted account features or professional suitability.

A controlled pilot

  1. Use a test workspace, public data or authorized synthetic material. Record the baseline and the selected product feature/plan.
  2. Run the primary case. Retain the actual prompt, output, source references, elapsed time, usage charges and action log.
  3. Run the altered-input/permission case. Inspect missing-data handling, escalation and unintended actions.
  4. Compare each observable pass condition with independent ground truth. Report failures and unresolved dependencies explicitly.
  5. Repeat only after a material product, configuration or test-input change. Obtain qualified review for professional-domain outputs.

Read the evidence and publication standard →

Executable coding test inputs

These repositories retain deliberate bugs and independent regression/boundary checks. Baseline and reference-patch runs validate the test input; the named coding products have not run them.

  • Cursor source and tests (ZIP) — 3 regression tests failed as expected; 5/5 boundary tests passed. Reference: 8/8 tests passed using an independently authored reference patch.
  • GitHub Copilot source and tests (ZIP) — 2 regression tests failed as expected; 5/5 boundary tests passed. Reference: 7/7 tests passed using an independently authored reference patch.
  • Cline source and tests (ZIP) — 9 regression tests failed as expected; 6/6 boundary tests passed. Reference: 15/15 tests passed using an independently authored reference patch.
  • Aider source and tests (ZIP) — 4 regression tests failed as expected; 5/5 boundary tests passed. Reference: 9/9 tests passed using an independently authored reference patch.

Product access attempts awaiting output

  • Tidio Lyro: The public demo did not load a chat input after selecting the ecommerce scenario and resetting it. No question was sent and no product reply was received. Checked 2026-10-02T04:55:08.477607+00:00.
  • Photoroom: The background-removal entry displayed a human-verification challenge. The marketplace demo was visible but its embedded controls could not be operated in this session; its independent entry requested an API key. No image was uploaded or processed. Checked 2026-10-02T04:55:08.477607+00:00.

An access limitation does not establish a product failure or a quality score.

Product evaluation cases

60 profiles · 120 planned cases · no invented ratings.

Download all controlled inputs and acceptance conditions (JSON) →

Download LangGraph local test evidence →

Download CrewAI local test evidence →

Download n8n local test evidence →

Recorded evaluation scope and planned product cases
Product and primary caseProduct-account testSource accessDirectory page
Shopify SidekickShopify store analysisAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
Helium AgentAmazon seller researchAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
KalodataTikTok Shop market researchAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
Gorgias AI AgentCommerce support resolutionAccess: accountNot executed2 specific planned cases3/3 accessible in last checkPassed
Tidio LyroKnowledge-based customer supportAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
PhotoroomProduct image preservationAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
Creatify AgentProduct advertisement draftAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
OpusClipLong recording to short clipsAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
DescriptTranscript-based editingAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
ElevenLabsAuthorized voiceover localizationAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
Buffer AI AssistantChannel-specific social copyAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
JasperBrand-guided campaign copyAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
n8nExplicit workflow orchestrationAccess: accountLocal runtime evidence recordedNot executed2 specific planned cases2/2 accessible in last checkPassed
Zapier AgentsHosted connected teammateAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
ElicitEvidence extraction from papersAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
HarveyPublic contract clause analysisAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
Lexis+ with ProtégéAuthority-backed legal researchAccess: accountNot executed2 specific planned cases0/1 accessible in last checkPassed
Patsnap EurekaDate-bounded prior-art retrievalAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
PlantixCrop photo screeningAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
AlphaSenseCited earnings researchAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
NansenRead-only wallet monitoringAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
ChainalysisTraceable blockchain investigationAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
RebuyCommerce recommendation workflowAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
HeyGenAvatar video localizationAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
OpenEvidenceMedical evidence questionAccess: accountNot executed2 specific planned cases1/1 accessible in last checkPassed
SpellbookContract drafting inside WordAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
Solve IntelligencePatent specification draftingAccess: sales demoNot executed2 specific planned cases1/1 accessible in last checkPassed
AGRIVI AI EngageFarm monitoring and advisoryAccess: sales demoNot executed2 specific planned cases2/2 accessible in last checkPassed
HebbiaMulti-document financial evidenceAccess: sales demoNot executed2 specific planned cases1/1 accessible in last checkPassed
DuneOn-chain SQL analysisAccess: accountNot executed2 specific planned cases4/4 accessible in last checkPassed
CursorRepository-scoped code changeAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
GitHub CopilotAssisted repository fixAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
Replit AgentRunnable prototype generationAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
ClinePermissioned coding workflowAccess: local installNot executed2 specific planned cases2/2 accessible in last checkPassed
OpenHandsSandbox repository repairAccess: local installNot executed2 specific planned cases1/1 accessible in last checkPassed
AiderTerminal code editingAccess: local installNot executed2 specific planned cases2/2 accessible in last checkPassed
Salesforce AgentforceCRM action governanceAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
HubSpot Agent HubCRM-grounded assistanceAccess: sales demoNot executed2 specific planned cases1/1 accessible in last checkPassed
LindyConnected operations assistantAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
ClaySource-linked prospect enrichmentAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
Fin by IntercomSupport knowledge and handoffAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
Zendesk AI AgentsService workflow evaluationAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
SciteCitation context assessmentAccess: accountNot executed2 specific planned cases0/2 accessible in last checkPassed
SciSpace Research AgentsPaper reading and extractionAccess: accountNot executed2 specific planned cases0/2 accessible in last checkPassed
KhanmigoGuided learning dialogueAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
NotebookLMSource-bounded knowledge synthesisAccess: accountNot executed2 specific planned cases4/4 accessible in last checkPassed
Microsoft Copilot StudioConfigurable enterprise agentAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
Glean AgentsPermission-aware enterprise searchAccess: sales demoNot executed2 specific planned cases2/2 accessible in last checkPassed
DustWorkspace assistant with actionsAccess: accountNot executed2 specific planned cases2/2 accessible in last checkPassed
DifyConfigurable agent applicationAccess: local installNot executed2 specific planned cases2/2 accessible in last checkPassed
CrewAIMulti-agent task coordinationAccess: local installLocal runtime evidence recordedNot executed2 specific planned cases2/2 accessible in last checkPassed
LangGraphStateful graph control flowAccess: local installLocal runtime evidence recordedNot executed2 specific planned cases2/2 accessible in last checkPassed
PikpopProduct concept and sourcing briefAccess: public demoNot executed2 specific planned cases0/1 accessible in last checkPassed
PostizDraft social content and scheduling reviewAccess: accountNot executed2 specific planned cases6/6 accessible in last checkPassed
TypefullySource-faithful social thread draftingAccess: accountNot executed2 specific planned cases6/6 accessible in last checkPassed
Taja AIYouTube metadata and source-context reviewAccess: accountNot executed2 specific planned cases3/3 accessible in last checkPassed
Photo AIAuthorized portrait identity and output reviewAccess: accountNot executed2 specific planned cases4/4 accessible in last checkPassed
Goblin ToolsActionable task decompositionAccess: public toolNot executed2 specific planned cases4/5 accessible in last checkPassed
Saner.AIPersonal knowledge retrieval with conflictsAccess: accountNot executed2 specific planned cases0/4 accessible in last checkPassed
SkyvernBrowser form automation with output approvalAccess: local installNot executed2 specific planned cases0/6 accessible in last checkPassed