Development

What a Legacy PHP Code Audit Should Deliver

A legacy PHP code audit should produce more than a list of old dependencies and style complaints. Its real deliverable is a defensible map from evidence to risk, then from risk to a sequence of practical actions.

The audit should help a team answer four questions:

  1. What does the system actually contain and do?
  2. Which conditions create meaningful business or operational risk?
  3. Which improvements should happen first?
  4. How can each recommendation be verified and reversed?

Agree on scope and access

Define the boundaries before reviewing code. A useful intake identifies:

  • Repositories, branches, and deployment artifacts in scope
  • Production and supported runtime versions
  • Framework, CMS, plugin, and dependency versions
  • Web, API, CLI, worker, and scheduled entry points
  • Databases, queues, caches, files, and external services
  • Authentication, authorization, payment, and privacy boundaries
  • Build, test, deployment, backup, and rollback procedures
  • Known incidents, support themes, and planned product changes

Read-only production telemetry can be valuable, but source access does not imply permission to access production data. Use sanitized samples and least-privilege credentials. Record what was unavailable because each gap limits the confidence of related conclusions.

An audit is not a penetration test unless that work is explicitly authorized and scoped. Passive secure-code review, dependency checks, and local reproduction are different from attacking a live system.

Build an evidence inventory

Capture facts in a form another engineer can reproduce. Typical evidence includes:

  • Output from php -v, php -m, and composer show
  • Composer platform constraints and audit results
  • Static-analysis configuration and findings
  • Test commands, coverage boundaries, and failures
  • Database schema, indexes, and representative query plans
  • Routes, hooks, jobs, and integration endpoints
  • Locations that read secrets or personal data
  • Deployment frequency, failed releases, and recovery steps
  • Dependency cycles and high-change modules

Use generated reports as inputs, not conclusions. A tool can identify an outdated package or a high-complexity function. The reviewer must still determine reachability, exploitability, business impact, and the cost of a safe change.

Composer's audit command checks installed packages against security-advisory data. It does not review custom authorization logic, insecure direct-object references, or operational secret handling. The OWASP secure code review guide provides a broader set of review concerns.

Review the system through several lenses

Runtime and dependency support

Identify unsupported PHP versions, unmaintained packages, abandoned extensions, pinned constraints, custom forks, and installation scripts that cannot be reproduced. Confirm findings against current official support information rather than relying on memory.

Security and privacy

Trace authentication and authorization separately. Review input validation, output encoding, database construction, file access, uploads, deserialization, session handling, secret storage, error disclosure, webhook verification, and outbound requests.

For every sensitive action, ask both “who may invoke it?” and “which object may that user act on?” A route can require login and still expose another customer's record.

Data integrity

Inspect transaction boundaries, retries, idempotency, unique constraints, migration practices, time zones, monetary values, and failure recovery. Determine which system owns each replicated field and how drift is detected.

Architecture and change risk

Map entry points to business capabilities and dependencies. Look for global mutable state, hidden service creation, circular dependencies, duplicated rules, and modules where unrelated changes repeatedly collide.

Do not call every large class a critical risk. Connect structure to an observable consequence such as untestable billing behavior, frequent regressions, or inability to upgrade a runtime.

Performance and operations

Use production-like evidence for slow queries, memory growth, cache effectiveness, job duration, external latency, and failure modes. Review monitoring and alerting as part of the application, not as an afterthought.

Delivery safety

Inspect how developers reproduce the environment, run tests, create releases, migrate data, deploy, and roll back. A modest code change can be high risk when releases are manual and restoration is untested.

Separate fact, inference, and recommendation

Every finding should make the reasoning visible:

  • Observed fact: Evidence directly verified in source, configuration, a test, or telemetry.
  • Inference: The risk that follows from the fact, including assumptions and confidence.
  • Recommendation: A proposed action, its intended outcome, and how to verify it.

This discipline prevents opinions from masquerading as facts. It also lets the team correct an assumption without discarding the useful evidence.

Here is a hypothetical example:

Fact: download.php accepts a numeric file ID, loads the file row, and streams its filesystem path. The route checks that a session is logged in, but does not compare the file's account ID with the session account. A local request using two fixture accounts returned the second account's file.

Inference: Any authenticated user who can discover or guess another valid file ID may retrieve that file. Confidence is high for the reviewed route. Other download paths were not tested.

Recommendation: Enforce account ownership in the database query, return the same not-found response for missing and unauthorized records, add cross-account integration tests, review all file-serving entry points, and monitor rejected access attempts without logging sensitive paths.

The sample includes evidence, scope, confidence, impact, and a testable correction. “Authorization needs improvement” does not.

Use a risk model the business can understand

Prioritize findings with explicit dimensions, for example:

  • Impact on confidentiality, integrity, availability, revenue, or compliance
  • Likelihood and required access
  • Exposure and reachability
  • Detectability and current controls
  • Blast radius
  • Remediation complexity and change risk

Use labels such as critical, high, medium, and low only after defining them. A public remote path to payment data and an internal code-duplication concern should not receive the same treatment because both are technically undesirable.

Add a separate confidence rating. A high-impact hypothesis with limited evidence may justify immediate validation, but it should not be presented as a confirmed vulnerability.

Deliver findings that are independently reproducible

Give every finding a stable identifier and enough evidence for another engineer to reproduce it safely. Keep the observed fact, inference, recommendation, affected capability, impact, scope limits, confidence, existing controls, dependencies, verification, and rollback in one record. The delivery checklist below is the single handoff inventory.

Remove credentials and personal data from screenshots, logs, and examples. Store sensitive security details in an access-controlled channel rather than a broadly shared slide deck.

Provide a roadmap, not a backlog dump

Group recommendations into outcome-oriented phases. A reasonable shape is:

Contain: Restrict exposed paths, rotate leaked secrets, disable dangerous behavior, or add monitoring where a credible immediate risk exists.

Stabilize: Make the environment reproducible, capture backups and rollback, add smoke tests, and document the critical system paths.

Modernize: Upgrade the runtime and dependencies in supported increments, improve type information, create boundaries, and remove duplicated business rules.

Optimize: Tune proven bottlenecks, simplify operational work, and remove temporary compatibility layers.

Do not assign universal durations before the team has validated dependencies and completed a representative slice. Provide ordering, prerequisites, decision points, and sizing assumptions. Let delivery estimates become more precise as uncertainty falls.

Use one delivery checklist

Different readers need different depth, but they should receive one connected package rather than several overlapping inventories. Use this checklist at handoff:

  • Scope, methods, repositories, environments, and access limitations
  • Executive summary of the major risks, immediate decisions, and confidence
  • System, data, dependency, and trust-boundary map
  • Prioritized finding records, each with its stable ID and complete evidence-to-action chain
  • Sequenced contain, stabilize, modernize, and optimize roadmap referencing those finding IDs
  • Engineering appendix with sanitized commands, outputs, query plans, and source references
  • Unanswered questions, missing evidence, assigned owners, and the next validation point

Write the executive summary as a view over the detailed records, not a separate source of truth. The roadmap should reference finding IDs, and each finding should link to its reproducible evidence. That structure lets leadership read the decisions while engineers can trace every statement without duplicated facts drifting apart.

Define success after the audit

An audit is useful only if findings can move into delivery. Assign owners for immediate containment, convert approved recommendations into bounded work, and schedule validation of high-impact unknowns.

Track outcomes such as supported runtime coverage, critical-flow tests, time to restore, eliminated public exposure, baseline static-analysis debt, and deployment reliability. Closing a ticket is not the same as reducing risk.

A strong legacy PHP code audit leaves the team with a shared factual model, not a fear-inducing report. It shows what is known, what remains uncertain, and which next step buys the most safety or delivery capacity.