AI agent research & evidence
Inspect original AI product inputs, outputs and test limits. Find short checks, native execution records and the method behind uAgentKit guides.
uAgentKit editorial compilation · operated by Steven Yang. AI-assisted research and condition checks; no independent testing laboratory or professional accreditation is claimed. Original results are retained even when they fail.
What each kind of evidence establishes
| Evidence | Useful for | Does not establish |
|---|---|---|
| Official source review | Documented features, conditions and vendor claims | Independent product performance |
| Quick product check | One first output against a short, fixed brief | General accuracy, typical latency or a product ranking |
| Native runtime or model record | The named version, configuration, input and observed result | All models, editions or production environments |
| Website acceptance check | Page rendering, links and evidence availability | A vendor product task being executed |
Use the quick-check method or inspect the existing full test plans and records. Conditions are assessed by the research agent unless a record explicitly identifies another reviewer. No aggregate quality score is computed.
Short checks: first output, visible limits
5 first-output records and 3 blocked access observations. These supplemental observations do not change the completion status of a product’s existing predefined cases. Each record states whether an output was obtained. Access blocks are not quality failures.
Buffer AI Assistant · passed for this brief
2026-10-07 · 1 submitted attempt · Tested entry
Draft a fictional refill-shop LinkedIn post while preserving a fixed-dollar offer.
Exact input and first output
Input
Write one LinkedIn post of no more than 60 words for a fictional refill shop. Facts: REFILL-15 gives USD 3 off an order of USD 20 or more; the offer ends 31 October 2026; it is not a percentage discount. Include all facts. Do not invent free shipping or claim this is 15% off. Draft only.
First output
Refill smarter: use code REFILL-15 to get USD 3 off an order of USD 20 or more (this is a USD 3 off, not a percentage discount). Offer ends 31 October 2026.
- passed: USD 3 off USD 20 minimum. Both values and the minimum order are retained.
- passed: 31 October 2026 deadline. The date is retained exactly.
- passed: No invented percentage discount or free shipping. It explicitly says the offer is not a percentage discount and adds no shipping promise.
- passed: 60 words or fewer. 32 whitespace-separated words.
Limits: One synthetic prompt and one first output; no reliability or accuracy estimate. Public free generator only; authenticated scheduling and publishing were not tested. Model, backend calls, token usage, request latency and actual cost were not measured. Condition review was performed by the research agent, not an independent human reviewer.
Goblin Tools · passed for this brief
2026-10-07 · 1 submitted attempt · Tested entry
Make a synthetic team message more professional while preserving dates and unapproved commitments.
Exact input and first output
Input
hey team, the demo is on 14 October 2026 at 10:30 UTC, not 13 October. send me your slides by 12 October. we have not approved any budget yet. please don't promise a launch date.
First output
大家好,关于演示安排,定于2026年10月14日10:30 UTC进行,非10月13日。请于10月12日前将您的演示文稿提交给我。当前尚未审批任何预算,请勿承诺正式上线日期。
- passed: Corrected demo date and time preserved. Output retains 14 October 2026 at 10:30 UTC and says it is not 13 October.
- passed: Slide deadline preserved. Output retains 12 October.
- passed: Unapproved budget and no launch-date promise preserved. Both constraints remain explicit.
Limits: One synthetic input and one first output. More professional mode, three chili icons; displayed version v3.14.1. The interface was Chinese and the English input produced Chinese output. This is not evidence of successful English-to-English rewriting. Output language was not a frozen pass condition. No email was sent; model, request latency, token usage and actual cost were not measured. Conditions were checked by the research agent, without independent human review or a statistical accuracy estimate.
ChatGPT · passed for this brief
2026-10-08 · 1 submitted attempt · Tested entry
Draft one English shop-support reply from a fixed synthetic order brief, without inventing missing details.
Exact input and first output
Input
For a synthetic shop-support test, write one English customer reply under 80 words using only these facts: customer Maya, order L-204 is packing, dispatch date is not confirmed, no tracking number is assigned, next update within two business days. Do not invent a delivery date, tracking number, discount, refund, or completed action. Return only the reply.
First output
Hi Maya, your order L-204 is currently being packed. The dispatch date hasn’t been confirmed, and no tracking number has been assigned. You’ll receive the next update within two business days.
- passed: English reply under 80 words. 31 whitespace-delimited words; English.
- passed: Preserve Maya, L-204 and packing. Maya, L-204 and being packed are present.
- passed: Explicitly preserve unknown dispatch date and no tracking number. The dispatch date remains unconfirmed and no tracking number is assigned.
- passed: Promise only an update within two business days; no invented delivery, tracking, discount, refund or completed action. Only the two-business-day update is promised. No delivery date, tracking number, discount, refund or completed action is invented.
Limits: One synthetic input and the first visible answer; no regeneration. UI displayed Pro; exact underlying model and serving version were not exposed. Existing signed-in account; account personalization and hidden settings were not isolated. The two products are separate observations, not a controlled model benchmark. No customer message was sent and no order system or connector was used. Latency, tokens and actual cost were not measured. Conditions were assessed by the research agent without independent human review; no overall accuracy or product-quality score is inferred.
Claude · passed for this brief
2026-10-08 · 1 submitted attempt · Tested entry
Draft one English shop-support reply from a fixed synthetic order brief, without inventing missing details.
Exact input and first output
Input
For a synthetic shop-support test, write one English customer reply under 80 words using only these facts: customer Maya, order L-204 is packing, dispatch date is not confirmed, no tracking number is assigned, next update within two business days. Do not invent a delivery date, tracking number, discount, refund, or completed action. Return only the reply.
First output
Hi Maya, Thanks for reaching out about order L-204. It's currently being packed. The dispatch date hasn't been confirmed yet, and no tracking number has been assigned so far. We'll send you an update within two business days. Best regards, Customer Support
- passed: English reply under 80 words. 42 whitespace-delimited words; English.
- passed: Preserve Maya, L-204 and packing. Maya, L-204 and being packed are present.
- passed: Explicitly preserve unknown dispatch date and no tracking number. The dispatch date remains unconfirmed and no tracking number is assigned.
- passed: Promise only an update within two business days; no invented delivery, tracking, discount, refund or completed action. Only the two-business-day update is promised. No delivery date, tracking number, discount, refund or completed action is invented.
Limits: One synthetic input and the first visible answer; no regeneration. UI displayed Opus 5.5 and extra-high in Chinese. The backend identity was not independently verified. Existing signed-in account; account personalization and hidden settings were not isolated. The two products are separate observations, not a controlled model benchmark. No customer message was sent and no order system or connector was used. Latency, tokens and actual cost were not measured. Conditions were assessed by the research agent without independent human review; no overall accuracy or product-quality score is inferred.
QuillBot · Blocked before output
2026-10-08 · 0 submitted attempts · Tested entry
Check access to one bounded product task; no generated output obtained.
Exact input and first output
Input
The workshop has moved to 18 November 2026 at 14:00 UTC, not 17 November. Please send your slides by 15 November. The budget has not been approved, so do not promise a launch date.
First output
No generated output was obtained.
- unverified: Obtain a first output for the prepared brief. The entered text appeared in Standard mode, but the UI continued to show zero words. Two button clicks did not yield an output; successful generation submission was not confirmed.
Limits: Access observation only; zero confirmed submitted generation attempts. This is not a product-quality failure. No new account, paid plan, deployed app or published presentation was created. Output quality, task accuracy, latency and actual cost remain untested.
Lovable · Blocked before output
2026-10-08 · 0 submitted attempts · Tested entry
Check access to one bounded product task; no generated output obtained.
Exact input and first output
Input
Planned, not submitted: Build a local-only reorder calculator. Inputs: daily sales 4, lead time 7 days, safety stock 5, stock 20. Expected reorder point 33 and gap 13. Reject negative inputs. No database, external connector or deployment.
First output
No generated output was obtained.
- unverified: Obtain a first output for the prepared brief. The access route displayed a login form with Google, GitHub, Apple, email and SSO options. No application-generation task was submitted.
Limits: Access observation only; zero confirmed submitted generation attempts. This is not a product-quality failure. No new account, paid plan, deployed app or published presentation was created. Output quality, task accuracy, latency and actual cost remain untested.
Gamma · Blocked before output
2026-10-08 · 0 submitted attempts · Tested entry
Check access to one bounded product task; no generated output obtained.
Exact input and first output
Input
Planned, not submitted: Make a three-slide English workshop brief from approved facts: workshop 18 November 2026, 14:00 UTC; slides due 15 November; budget unapproved. Keep the facts and show budget as unknown. Do not publish.
First output
No generated output was obtained.
- unverified: Obtain a first output for the prepared brief. The free-start route displayed signup. No deck-generation task was submitted.
Limits: Access observation only; zero confirmed submitted generation attempts. This is not a product-quality failure. No new account, paid plan, deployed app or published presentation was created. Output quality, task accuracy, latency and actual cost remain untested.
Google SynthID Detector · passed for this brief
2026-10-08 · 1 submitted attempt · Tested entry
Record one supported-watermark check on a Google-hosted showcase JPEG, preserving the exact result and its scope.
Exact input and first output
Input
One unmodified downloaded Google Gemini Image showcase JPEG; 172827 bytes. SHA-256: 129c7d01a75813f197a018c6025861fb9d48bdb9bf29d59efc06342cafe56df0. Source page: https://deepmind.google/models/gemini-image/ This is a web-distributed showcase, not a newly generated original model export.
First output
SynthID was detected This media was made or edited with Google AI Please note that this media may have been edited further since then.
- passed: Retain exact file bytes, source URL and hash. The downloaded file, source asset URL, 172,827-byte size and SHA-256 hash were retained before submission.
- passed: Obtain and preserve the exact portal result or error. The first result reported a detected SynthID watermark and named Google AI. Its further-editing warning is retained.
- passed: Do not infer universal AI detection, accuracy percentage, human authorship or commercial clearance. The record reports one supported-watermark finding only. No general authorship, accuracy percentage, product-quality or rights conclusion is inferred.
Limits: One UI upload in the available signed-in session, without regeneration or adversarial modification. Official Google web showcase asset; original generation timestamp, exact producing configuration and private export were unavailable. The source CDN may have resized or compressed the image. No original camera-photo control, video, audio or text detection was tested. No accuracy or false-positive rate can be estimated from this sample. Quota, actual cost and latency were not measured. A missing charge display does not establish free or unlimited access. Conditions assessed by the research agent; independent human review is not claimed. The screenshot shows the actual result, not a vendor UI illustration.
Products with an executed native record
41 catalog products have at least one recorded execution. The list includes failed and partial results, and mixes model, public UI and control-flow evidence. It is ordered by the directory, not quality. Open a profile to see the configuration, case outcomes and original downloads.
- n8n: cases and evidence
- Cline: cases and evidence
- OpenHands: cases and evidence
- Aider: cases and evidence
- NotebookLM: cases and evidence
- CrewAI: cases and evidence
- LangGraph: cases and evidence
- Pikpop: cases and evidence
- Skyvern: cases and evidence
- AnythingLLM: cases and evidence
- Open WebUI: cases and evidence
- Open Interpreter: cases and evidence
- Rembg: cases and evidence
- Khoj: cases and evidence
- Vibe: cases and evidence
- LLM: cases and evidence
- LibreTranslate: cases and evidence
- Kokoro Web: cases and evidence
- Chatbox Community Edition: cases and evidence
- SiYuan: cases and evidence
- SillyTavern: cases and evidence
- big-AGI: cases and evidence
- Subtitle Edit: cases and evidence
- Fabric: cases and evidence
- GPT Researcher: cases and evidence
- AIChat: cases and evidence
- ShellGPT: cases and evidence
- tgpt: cases and evidence
- Mods: cases and evidence
- Chatblade: cases and evidence
- gptme: cases and evidence
- GPTScript: cases and evidence
- OpenCode: cases and evidence
- Goose: cases and evidence
- Crush: cases and evidence
- Pi: cases and evidence
- nanobot: cases and evidence
- PicoClaw: cases and evidence
- Lovart: cases and evidence
- Kling AI: cases and evidence
- Patsnap Patentability Grader: cases and evidence
Responsibility, corrections and reuse
The site operator is Steven Yang. Compilation and testing assistance are disclosed; named expert or independent peer review is not claimed. See the editorial policy and commercial disclosure. To correct a record, provide its URL, the disputed statement and a supporting source through the correction form.
When citing a result, include the product, run date, configuration, input, outcome and evidence URL. Downloads may contain third-party material under its own terms; check those terms before reuse. The existence of a record or citation does not imply vendor endorsement. External recognition and Google search performance are separate from the tests shown here.