karawaci.kode

2026-06-25 · 7 min

Capacity Planning Payment Gateway Tangsel: 0-10k Tx/Hari

Delapan bulan lalu, payment gateway klien saya di Tangerang Selatan (B2B aggregator untuk merchant kecil) baru launch dengan 200 transaksi/hari. Bulan lalu mencapai 10,000 transaksi/hari. Saya share capacity planning yang berhasil tanpa over-spend, plus 2 momen saya hampir gagal capacity.

Baseline awal

Stack lama (yang saya inherit):

  • Bun 1.3 + Hono di Hetzner cx21 (2 vCPU, 4GB RAM)
  • Postgres 17 di same VPS (anti-pattern, will fix)
  • Redis di same VPS (anti-pattern)
  • Single instance, no LB

Throughput awal:

  • 200 tx/hari = ~0,02 tx/sec average
  • Peak (lunch + dinner): ~3 tx/sec
  • P95 latency: 240ms
  • Resource usage: 12% CPU, 1,8GB RAM

Way over-provisioned untuk launch. Sengaja: klien bayar saya untuk “build it right from start”.

Growth curve

Tx/hari per bulan:

  • Bulan 1: 200
  • Bulan 2: 480
  • Bulan 3: 1,100
  • Bulan 4: 2,400
  • Bulan 5: 4,200
  • Bulan 6: 6,800
  • Bulan 7: 8,500
  • Bulan 8: 10,200

Growth ~70%/bulan rata-rata. Acquisition driver: marketing partner-merchant (warung, toko sembako, jasa servis HP).

Capacity planning framework

Saya pakai pendekatan:

  1. Identify peak-to-average ratio (history data)
  2. Plan untuk 8x daily average di peak hour (safety margin)
  3. Threshold: scale up saat sustained P99 > 500ms ATAU error rate > 0,1%
  4. Threshold: scale down saat resource utilization < 30% selama 7 hari berturut-turut

Pattern ini lazy (reactive), tapi works untuk growth rate predictable.

Scaling events

Event 1: Bulan 3 (1100 tx/hari)

Pertama kali saya scale up. Trigger: P99 latency naik dari 280ms ke 620ms saat peak.

Action:

  • Separate Postgres ke VPS sendiri (cx31, $15/mo)
  • Tambah connection pooler (PgBouncer)
  • Original app tetap di cx21

Cost: +$15/mo. Time to implement: 6 jam.

P99 turun ke 320ms. Headroom buat 2 bulan ke depan.

Event 2: Bulan 5 (4200 tx/hari)

P99 spike lagi di peak (sekitar 1100 tx pernah dalam 1 jam). Root cause analysis: Redis di same VPS dengan app, kena memory pressure saat cache hit pattern berubah.

Action:

  • Move Redis ke VPS sendiri (cx21, $5/mo Redis-dedicated)
  • Upgrade app VPS dari cx21 ke cx31 (4 vCPU, 8GB RAM): $15/mo

Cost: +$15/mo (cx21 retire, cx31 baru) + $5/mo (Redis VPS). Net: +$20/mo.

Time to implement: 8 jam termasuk Redis migration tanpa downtime (slave-of pattern).

Event 3: Bulan 6 (6800 tx/hari) — saya hampir gagal

Maintenance Sabtu pagi rutin. Saya update Bun 1.3.245 → 1.3.247 (patch fix). Aplikasi restart 30 detik. Tidak ada announce.

Sabtu pagi 10:30 — peak time untuk top-up e-wallet merchant. Selama 30 detik restart: ~85 transaksi failed (queue tidak retry, langsung error ke client).

Klien dapat 14 complaint dari merchant dalam 1 jam. Saya panik fix. Patch keluar bahwa retry mechanism harus added di middleware.

Implementation post-mortem:

  • Idempotency key untuk semua write (klien generate UUID per attempt)
  • Retry-Safe pattern: 502/503/504 dari my server → client retry up to 3x dengan exponential backoff
  • Zero-downtime deploy: pakai systemd socket activation + drain pattern (lihat systemd VPS deploy saya)

Damage: small reputation hit, klien deduct Rp 4jt dari invoice bulan itu. Painful lesson.

Event 4: Bulan 7 (8500 tx/hari) — DB bottleneck

P99 naik bertahap. Hitung: 8500/hari peak 8x = ~1100/jam → ~0,3 tx/sec average, peak ~6 tx/sec. App fine. DB connection pool saturated di peak.

PgBouncer default max_client_conn = 100. Apps spawn lots of short connection. I raised:

  • max_client_conn = 500
  • default_pool_size = 25
  • Plus tweak pool_mode = transaction (faster than session mode)

P99 turun back to 280ms. Free fix, no infra cost.

Event 5: Bulan 8 (10,000 tx/hari) — DB IO

Postgres EBS-equivalent di Hetzner sudah hit IOPS ceiling untuk WAL write saat peak. Saya:

  • Migrate Postgres ke Hetzner CCX (dedicated CPU + faster local NVMe). cx31 → ccx13: $25/mo upgrade.
  • Tune wal_buffers, synchronous_commit = local untuk non-critical write.

Cost: +$10/mo (ccx13 vs cx31 diff).

P95 latency stable 220ms, P99 380ms di peak. Acceptable.

Total cost evolution

  • Bulan 1: $5/mo (cx21 all-in-one)
  • Bulan 3: $20/mo (app + DB terpisah)
  • Bulan 5: $40/mo (+ Redis VPS + app upgrade)
  • Bulan 8: $50/mo (+ Postgres upgrade)

Per-transaction cost (infrastructure):

  • Bulan 1: $5 / (200 × 30) = $0,00083 per tx
  • Bulan 8: $50 / (10,000 × 30) = $0,00017 per tx

5x lebih efficient per-tx at scale. Standard scaling economics.

Monitoring stack

Saya monitor (Prometheus + Grafana di Hetzner cx21):

  • payment_request_duration_seconds{endpoint, status} (histogram)
  • payment_request_total{endpoint, status} (counter)
  • db_connection_pool_size{state} (gauge)
  • redis_cache_hit_ratio (gauge)
  • bun_memory_rss_bytes (gauge)
  • External: Midtrans/BCA API latency

Alert (Alertmanager → Telegram):

  • P99 > 800ms sustained 5 menit → warn
  • Error rate > 0,5% sustained 2 menit → warn
  • Error rate > 1% sustained 1 menit → page

Alert page saya: dial Telegram + SMS (Twilio). 2 page dalam 8 bulan. Both legitimate, both saya respond < 5 menit.

Yang saya rekomendasi untuk capacity planning SMB

  1. Track history dari awal. Bahkan saat traffic kecil. Pattern muncul saat data 3-6 bulan terkumpul.

  2. Plan untuk 8x peak, bukan 2x average. Indonesia traffic spike di lunch hour + gajian akhir bulan.

  3. Scale komponen yang bottleneck, bukan blanket upgrade. App, DB, Redis bisa beda growth rate.

  4. Idempotency dari hari 1. Retry pattern. Worth setup awal, painful retrofit.

  5. Zero-downtime deploy. Pakai pattern systemd socket activation atau LB drain. Restart tanpa down impact.

  6. Tidak over-provision. Saya dulu pernah klien lain pakai cx51 untuk 50 tx/hari “biar safe”. Buang uang $30/mo selama 6 bulan = Rp 2,8jt sia-sia.

Verdict

Capacity planning SMB payment Indonesia: reactive scaling OK untuk growth predictable. Proactive monitoring + alert wajib. Right-size resource per komponen.

Bukan magic — saya gagal 1x (zero-downtime deploy lupa) yang kena reputation. Lessons learned, pattern stick.

Pattern saya yang reusable: monitoring stack + 8x-peak rule + per-component scaling. Sekarang saya pakai di 3 klien payment / fintech-adjacent. Cost-effective tanpa over-engineering.

Ditulis oleh Reza Pradipta