Priced by rows generated, not by data shared.

No per-seat fees. No data upload required for schema-defined generation. Pay for synthetic output — nothing else.

50K
free rows per month to start
5M
rows/month on Build tier
<4ms
per-row generation at batch scale (tabular, 20-column schema)

One plan for every stage of your project.

Prototype
Explore the pipeline
$0 /month
50,000 synthetic rows/month
  • Up to 20 columns per schema
  • CSV + JSON Lines output
  • Basic distribution matching
  • Community support
  • API access (rate-limited)
Most popular
Build
For active ML teams
$149 /month
5M synthetic rows/month
  • Unlimited columns per schema
  • CSV, Parquet, JSON Lines, S3 push
  • Advanced tail-weighting controls
  • Constraint solver (referential integrity)
  • Fidelity + utility quality reports
  • Email support (48h response)
Scale
For production pipelines
$499 /month
50M synthetic rows/month
  • Everything in Build
  • Dedicated generation workers
  • Custom constraint schemas
  • Privacy score reports (NN distance)
  • SLA: 99.9% uptime on generation API
  • Priority support + engineering calls

Annual billing available: Build at $119/mo, Scale at $399/mo. All prices in USD.

Questions we get from ML and data engineering teams.

No. When you upload a sample CSV, Twynvex fits a statistical model (marginal distributions + pairwise correlation matrix) and does not retain the raw rows. The model parameters are stored in association with your schema object; the raw records are discarded after fitting. Schema-defined generation (no sample at all) never touches real data by design.
Generation jobs that would exceed your monthly row limit are queued and held until the next billing cycle resets your quota, or until you upgrade your plan. You won't be billed overage automatically. Prototype tier users can upgrade to Build at any time mid-month — your new quota takes effect immediately.
Twynvex is a privacy-by-construction tool — no real patient records enter or exit the generation engine. Whether the output qualifies as de-identified or non-PHI under your organization's specific HIPAA interpretation is a legal determination your team makes with your legal counsel. We describe the technical architecture; we don't certify compliance outcomes. Many healthcare ML teams use Twynvex for model development precisely because it operates from schema definitions rather than real records.
Fidelity is reported as Jensen–Shannon divergence between the marginal distribution of each column in the real sample and the corresponding column in the synthetic output. JS divergence of 0 means identical distributions; 1 means completely different. Twynvex reports per-column JS and an aggregate weighted average. A utility score (train-on-synthetic, test-on-real AUC) is also provided as a practical measure of whether the synthetic data is usable for training.
Yes. Teams generating more than 50M rows/month, or requiring on-premises deployment, dedicated infrastructure, or custom SLAs, should contact us directly. We'll discuss volume pricing and infrastructure options based on your actual generation workload.

Start generating in 5 minutes.