A vendor’s penetration test is genuinely manual when human reasoning materially changes what is tested, how weaknesses are combined, how impact is validated and how evidence is explained. Manual work is not a person clicking “run” in a tool, nor does it mean tools are absent. The proof is observable in the role matrix, product-specific hypotheses, tailored demonstrations, validation trail and live technical debrief.

A CTO can evaluate this before purchase by asking the proposed delivery team to explain its workflow against one real authorization boundary and one business process. Marketing labels are easy; a coherent test design is harder to fake.

Manual work begins before payloads

The tester learns the product’s actors, objects, actions, states and trust boundaries. It identifies prohibited outcomes and develops hypotheses. For a multi-tenant workflow, that means understanding which tenant owns the object, which roles may act, how membership changes and where enforcement occurs.

  • Attack-surface map reconciled with architecture and observed client traffic.
  • Role, tenant and permission matrix.
  • Critical workflow and state-transition model.
  • Product-specific abuse cases and test priorities.
  • Explicit sampling and coverage limits.

Manual authorization evidence

Ask how the team compares users, roles and tenants. A strong answer describes paired accounts, controlled objects, request replay, server-side decisions and secondary effects. It does not reduce authorization testing to changing one numeric identifier.

The final evidence should identify both identities, the expected rule, the request or action that crossed it and the unauthorized result. For sensitive systems, proof should minimize exposure rather than browse real data.

Manual business-logic evidence

Look for test cases involving sequence, timing, quantity, state and channel. Examples include skipping approval, replaying a credit, racing an allocation, acting after cancellation, using a mobile-only endpoint from a web token or triggering a background job with altered tenant context.

The provider should explain how it learns the intended rule and how it distinguishes a security abuse from unusual but valid functionality.

Human validation and chaining

Scanners generate observations. A tester reproduces them, removes false positives, determines preconditions and explores whether several weaknesses create a material attack path. The report should separate the original signal from the validated finding.

  1. Discover a signal through tooling, review or manual exploration.
  2. Reproduce it consistently under controlled conditions.
  3. Identify identity, state and environmental preconditions.
  4. Explore safe variants and relevant chains.
  5. Stop at the minimum evidence that establishes impact.
  6. Have a reviewer check the conclusion and remediation.

What to request before selection

  • A redacted report containing requests, responses, identity context and root-cause remediation.
  • A sample role-and-tenant coverage matrix.
  • An explanation of tester and reviewer responsibilities.
  • A methodology showing where tools stop and human analysis begins.
  • A description of critical-finding escalation and safe proof standards.
  • A live discussion with the proposed technical lead.

Questions that expose shallow claims

  • Which hypotheses would you create from our tenant and role model?
  • How do you test workflow sequence and state transitions?
  • How do you discover APIs the UI does not document?
  • How do you validate scanner observations and consolidate duplicates?
  • How do you decide when exploitation has gone far enough?
  • How will untested combinations appear in the report?

Tool use is not a disqualifier

Good testers use proxies, crawlers, fuzzers, scripts and scanners for repeatability and breadth. They may automate comparisons across many objects or roles. The manual value is in designing the comparison, interpreting differences, choosing safe follow-up and explaining impact. A claim of “100% manual” can itself be a warning when it implies weak discovery or no repeatable regression support.

Delivery signals during the engagement

  • Questions become more specific as testers learn the product.
  • Urgent findings contain validated evidence rather than raw alerts.
  • Testers explore alternate roles, endpoints and workflow sequences.
  • Status updates discuss coverage and blockers, not only finding counts.
  • The technical debrief explains attack paths and root causes without reading slides.

Red flags

  • The scope and duration are fixed before roles or workflows are known.
  • The methodology is primarily a list of tool names.
  • Every finding matches generic scanner language and screenshots.
  • The provider promises complete coverage or no false positives in advance.
  • The proposed tester is unavailable for technical evaluation or debrief.
  • Retesting consists only of rerunning the original scanner.

Put manual depth in the SOW

Require attack-surface discovery, relevant automated support, manual authorization and business-logic testing, exploit validation, technical review, a coverage record and developer-ready evidence. Protect the hands-on window from access delays and scope expansion. Avoid a vague percentage of “manual effort” that cannot be verified.

The OWASP Web Security Testing Guide offers useful testing areas. Genuine manual depth appears when the provider converts applicable areas into product-specific roles, workflows, hypotheses and evidence.

The practical decision

Believe observable engineering work, not the label. A genuine manual penetration test shows how humans model the product, compare identities, manipulate state, validate impact, build attack paths and explain root cause—while using automation where it improves reach and repeatability.

Ask WIMD to demonstrate the manual test approach for your scope using one real role boundary and workflow.