Two VAPT proposals that both claim “OWASP coverage” may be selling materially different work. OWASP is a family of valuable security resources, not a standard-sized package, fixed test duration or proof of manual depth. Compare proposals by normalizing scope, test method, people, evidence, reporting, remediation and retesting before comparing price.
The cheaper proposal may be efficient and appropriate. It may also omit roles, APIs, business-logic testing, quality review or closure support. The more expensive proposal may provide deeper assurance, or it may simply contain more contingency and branding. A disciplined comparison makes those differences visible instead of using the OWASP label as a substitute for detail.
First decode the “OWASP coverage” claim
Ask which resource and version the provider means. The OWASP Web Security Testing Guide contains broad testing areas; the OWASP Top 10 is an awareness document; the ASVS is a verification framework; and OWASP also publishes API and mobile resources. A proposal should explain how relevant risks become test activities for your product.
A statement such as “covers OWASP Top 10” does not tell you which authenticated roles are compared, whether tenant isolation is tested, how APIs are discovered, whether business-logic abuse is explored, or how findings are validated. Treat it as a starting vocabulary, not as an acceptance criterion.
Build one comparison scope
- List every application, domain, API, mobile build, environment and integration boundary.
- List user roles, privileged identities, support paths and tenant combinations.
- Identify high-impact workflows such as approvals, payments, exports, sharing and account recovery.
- State production restrictions, testing hours, prohibited actions and data constraints.
- Define the required report audiences, evidence quality, retest and closure artifact.
Send the same baseline to each provider and require them to mark included, excluded, assumed and optional items. A proposal that cannot be mapped to the baseline is not ready for commercial comparison.
Score the proposals across eight dimensions
1. Scope completeness
Check that all intended surfaces, roles and workflows appear. Look for limits hidden in footnotes: one role only, one API collection, no mobile backend, no production validation or a fixed endpoint count that does not match the architecture.
2. Manual testing depth
Ask how testers will examine authorization, tenant separation, business logic, workflow state and chained weaknesses. Tooling should support discovery and repeatability; it should not be the entire method. The provider should explain product-specific hypotheses, not only list scanners.
3. Tester capability and review
Identify who will perform the work, who will review it and what relevant application or API experience they have. Certifications can support confidence but do not replace named responsibilities, sample evidence and a credible quality process.
4. Schedule realism
Compare readiness, protected testing time, serious-finding notification, draft, factual review, final report and retest. One vendor’s “five days” may mean hands-on testing while another’s means the full calendar. Normalize the milestones.
5. Finding evidence
Review a redacted sample. Findings should include affected assets, preconditions, reproduction, proof, impact, severity reasoning and actionable remediation. A scanner screenshot and generic definition are not equivalent to validated evidence.
6. Reporting for each audience
Leadership needs exposure, priorities and decisions. Engineers need requests, responses, identities, root cause and repair guidance. Customers or auditors may need scope, dates, method and closure status without sensitive exploit detail. Confirm which outputs are included.
7. Remediation collaboration
Ask whether testers will explain findings to engineers, review proposed fixes and identify systemic causes. Define response times and the duration of clarification support. “Support included” is too vague to evaluate.
8. Retest and closure
Record the number of cycles, eligible findings, scheduling conditions and final artifact. A rerun of the original scanner is not sufficient when the finding involved authorization or workflow logic.
Use a weighted score instead of one total
Weight dimensions according to the buying event. An enterprise customer deadline may emphasize accepted evidence and turnaround. A critical product launch may emphasize role coverage, manual depth and production safety. A remediation-heavy retest may emphasize continuity with the original tester and closure quality.
Score each proposal only where the vendor supplied evidence. Mark uncertainty rather than awarding points for polished language. Then review the largest score differences with the vendors. The purpose is not to create a mathematically perfect winner; it is to expose material assumptions before purchase.
Commercial items to normalize
- Taxes, travel, onboarding and any platform or project-management fees.
- Named tester effort, parallel staffing and technical-review effort.
- Draft and final reports, executive readout and engineering walkthrough.
- Retest cycles, finding limits and expiration period.
- Change requests for new endpoints, roles, builds or environments.
- Cancellation, access-delay and rescheduling rules.
- Evidence retention, secure transfer and deletion obligations.
A low total may become expensive if retesting, additional roles and reporting are separate. A higher total may include work you do not need. Compare the cost of the normalized required outcome, not the front-page number.
Questions that reveal whether the offers are equivalent
- Show how the proposed days or effort map to discovery, manual testing, validation, reporting and review.
- Which roles and tenant pairs will be tested against which actions?
- How will undocumented or client-discovered APIs be handled?
- What is excluded from the OWASP statement?
- What must be ready before the delivery clock begins?
- Who decides whether a finding is valid and how disagreements are recorded?
- What exactly happens after engineers submit a fix?
Red flags in either proposal
- A fixed quote was issued without asking about identities, APIs or workflows.
- “100% OWASP coverage” is promised without defining applicable tests or limitations.
- The provider guarantees no false positives or a clean certificate before testing.
- Every application receives the same duration regardless of complexity.
- The report sample contains raw tool output and generic remediation.
- The vendor refuses to name exclusions, tester roles or retest terms.
The practical decision
Select the proposal whose coverage, method, people and deliverables best support the business decision at a proportionate cost. OWASP alignment is useful only when translated into product-specific testing. If one quotation is much cheaper, identify which normalized rows explain the difference; do not assume identical assurance from an identical label.
Use WIMD’s penetration-testing cost and scope overview as a starting point, then ask every shortlisted provider to respond against the same role, workflow and evidence matrix.
