Paid engagement

AI Reliability Evaluation

A fixed-scope, fixed-price look at what your GenAI application's telemetry actually shows: where operations fail, which provider and model carry the errors and the latency, and which evidence is missing rather than inferred.

This is an evaluation, not a pilot and not a production deployment. The boundary below is part of the offer, not fine print.

What you get

Fixed at engagement start. Anything outside this is re-scoped and requoted, not absorbed.

Price
USD $7,500 per AI Reliability Evaluation
Coverage
One application or environment
Data ceiling
Up to 100 GB of approved telemetry
Window
A 10-business-day window
Delivery effort
At most 30 delivery hours
Workspace
One workspace
Output
One versioned assessment and one readout
Wind-down
Export and deletion evidence at engagement end

What the evaluation answers

A bounded question set, fixed in advance. Each query carries a stable identifier, a budget, and a result fingerprint, so a finding can be re-run and checked rather than taken on trust.

  • Operation volume, error rate, and latency by provider and model.
  • Input and output token use, and first-chunk latency, where your telemetry supplies them.
  • Prompt and config version comparisons, without storing prompt bodies by default.
  • Tool and retrieval failure candidates, with exact trace pivots.
  • Affected sessions or agents, where tenant-safe identifiers exist.
  • Missing or redacted evidence, stated as missing — never inferred content.

Fixture self-assessment

Inspect the sample before you schedule

We ran the frozen question pack through our shipped GenAI fixture and published the resulting sample assessment, before/after report, evidence-gap audit, and workspace captures. This is synthetic fixture material from our own self-assessment — not customer data, a live run, payment evidence, or A2-M2 exit evidence.

first-party evaluation rehearsal — not A2-M2 evidence

Where the boundary sits

Stated up front so you can tell whether this engagement fits before you pay for it.

PackDB is not your system of record

The input is a customer-approved export, a sanitized fixture, or an explicitly non-authoritative shadow copy. Your source telemetry stays in your authoritative system throughout.

No production SLO

This engagement carries no production availability, durability, recovery, or migration commitment. Continuous authoritative ingestion or a dependency on this workspace for production alerting is outside the evaluation boundary.

Fixed duration, ceiling, and owner

The engagement has a fixed duration, a data-volume ceiling, a named owner, and a documented support channel — all agreed before it starts.

Privacy enforced before the write

The packdb.genai.v2 profile enforces privacy before the write-ahead log. Every report identifies the policy, profile, source window, completeness, and redactions. No prompt, response, tool, or retrieval body appears in any report.

Your artifacts are exportable, and the teardown is documented

You receive exportable assessment artifacts and a documented deletion and teardown result at the end, including any residual backend-retention interval where immediate physical erasure is unavailable.

Deterministic findings

The workspace is deterministic. An analyst may write findings, and any generated hypothesis is labeled as analyst or model interpretation — it never replaces the query result underneath it.

Scope an evaluation with us

Tell us the application, roughly how much telemetry it produces, and what you need answered. We will confirm the fit against the boundary above before quoting.

Schedule AI Evaluation