A serious penetration test for an application with multiple roles, tenants and APIs cannot be estimated credibly from page count alone. Duration follows the number and complexity of trust relationships the tester must exercise: identity pairs, tenant boundaries, workflows, interfaces, platforms, states and integrations. A three-day quote may be appropriate for a tightly bounded component, but it is a warning sign when it claims deep coverage of a broad product without an explicit model.
Ask vendors to show how they derived the effort. The purpose is not to maximize days. It is to connect time to meaningful coverage, surface assumptions and decide what can be tested deeply within the available window.
Build the estimate from test dimensions
Identity and role combinations
Each role adds more than one login. Testers need to compare what an actor may do with what peers, higher roles, lower roles and unauthenticated users must not do. Lifecycle states such as invited, suspended, deactivated or recently reassigned create additional authorization conditions.
- Number of materially different human roles.
- Machine identities, API clients and service accounts.
- Horizontal peer relationships within the same role.
- Role transitions, delegation, impersonation and recovery.
- Authentication methods, session types and token lifecycles.
Tenant relationships
Multi-tenancy creates an isolation matrix. The tester needs controlled identities and data in at least two tenants, plus an understanding of shared services. Parent-child organizations, cross-tenant sharing, support operations and data migration increase the number of meaningful paths.
- Tenant creation, membership and ownership changes.
- Object identifiers and APIs that carry tenant context.
- Exports, searches, caches, queues and background jobs.
- Administrative or support access across customer boundaries.
API and client surface
Count distinct operations and authorization behaviours, not merely documented paths. Versioned APIs, GraphQL, WebSockets, mobile backends, webhooks, file processors and internal interfaces may require different tooling, credentials and state setup.
Workflow depth
Bespoke business logic requires time for exploration and sequence changes. Approval, payment, entitlement, onboarding, recovery, import and destructive workflows should be tested across relevant roles and states—not just as isolated requests.
Account for the full engagement lifecycle
- Scoping and architecture familiarization.
- Environment validation, account setup and data preparation.
- Discovery and automated assistance where appropriate.
- Manual testing of access control, business logic and abuse paths.
- Evidence review, reproduction and severity analysis.
- Report writing and technical quality assurance.
- Readout, remediation clarification and a later retest.
A quote that mentions only “testing days” can hide whether reporting, project management, setup and retesting are included. Ask for calendar time and practitioner effort separately.
Environment readiness changes effective duration
Testing time is easily consumed by missing credentials, unstable builds, incomplete seed data, blocked source addresses and undocumented request requirements. Determine whether the estimate assumes a ready environment and how blocked time will be handled.
- Representative build and production-like configuration.
- Working accounts for every agreed role and tenant relationship.
- API specifications, request examples and integration credentials.
- Known feature flags and environmental differences.
- Fast engineering support for access and data restoration.
Readiness should not be used to compress security analysis. Complete setup before the testing window where possible and hold an early access validation session.
Ask how coverage will be sampled
Not every permutation deserves equal depth. A capable tester groups equivalent controls, samples representative paths and spends more time where business impact or architectural novelty is highest. The proposal should explain that selection.
For example, endpoints using one shared authorization policy may be tested through representative operations plus attempts to bypass the common enforcement point. A bespoke support-impersonation workflow deserves independent analysis even if it exposes only a few routes.
What a credible effort model looks like
- A base amount for architecture review, setup, reporting and quality assurance.
- Added effort for distinct identity and tenant relationship families.
- Added effort for API breadth, protocols and client platforms.
- Added effort for critical bespoke workflows and integrations.
- Explicit assumptions about environment readiness and exclusions.
- Separate time for clarification, remediation support and retesting.
The model need not expose a vendor’s internal commercial formula. It should be detailed enough for engineering to see what gains or loses coverage when scope changes.
Evaluate a three-day quote carefully
Three days can support a focused test of a small service, a narrow release change or a small set of critical workflows. It is unlikely to support comprehensive manual coverage of a feature-rich multi-role, multi-tenant application and its APIs once setup, evidence and reporting are included.
Questions to ask
- How many practitioner days are testing versus reporting?
- Which roles, tenant pairs, clients and API versions are included?
- Which critical workflows receive manual business-logic testing?
- What is sampled, and what is explicitly excluded?
- Does the estimate include retesting and remediation clarification?
- What happens if discovery reveals additional in-scope interfaces?
A short engagement is not automatically poor. An unexplained promise of broad coverage is the concern.
Choose between more time and narrower claims
When the budget or release window is fixed, reduce the scope deliberately. Prioritize exposed and high-impact workflows, core identity and tenant boundaries, and the interfaces that perform sensitive actions. Record deferred surfaces and plan the next increment.
Do not keep the broad scope label while silently reducing depth. Stakeholders must be able to distinguish a focused risk-based assessment from comprehensive application coverage.
Plan calendar time around engineering decisions
Allow time before testing for scope resolution and account validation, during testing for access support and emerging findings, after testing for report QA and clarification, and after remediation for retesting. Calendar duration is often longer than hands-on testing effort because fixes and evidence cross team schedules.
The methodology in NIST SP 800-115 is a useful reminder that planning, execution, analysis and reporting are all parts of a technical security assessment—not optional overhead around a scan.
Compare vendors using coverage assumptions
Normalize proposals before comparing price or duration. Place each vendor’s included assets, roles, tenant relationships, workflows, API protocols, reporting depth and retest terms side by side. A lower quote may reflect automation, experience or a narrower interpretation; only the assumptions reveal which.
The practical decision
Estimate a multi-role, multi-tenant pentest from identities, isolation relationships, APIs, workflow depth, platforms, integrations and environment readiness. Require explicit sampling and exclusions. If the available time cannot support the desired depth, narrow the claim and schedule further coverage rather than accepting an opaque assurance.
Ask WIMD for an evidence-based pentest effort estimate tied to your actual roles, tenants, APIs and critical workflows.
