, ,

How to Audit Journal Entries: The Management Override Lens, With a Worked Year-End Run

In May 2002, an internal audit team at WorldCom led by Cynthia Cooper started pulling apart journal entries that moved ordinary network expenses — “line costs,” the fees paid to other carriers — out of the income statement and into capital accounts. The entries were large, round, recorded after the close, and unsupported. By June 25 the company had disclosed more than $3.8 billion of improperly capitalized line costs ($3.055 billion for 2001 and $797 million for Q1 2002); by August, reserve manipulations had roughly doubled the total; the final accounting irregularities reached roughly $11 billion, and the company’s collapse became the largest bankruptcy in US history to that point. Every control WorldCom had — approval matrices, reconciliations, an external audit — was defeated by the one mechanism no process control constrains: senior management writing entries directly into the general ledger.

That is why journal entry testing exists, and why it is unlike every other test in your program. Most controls testing asks “does the process work?” JE testing asks the darker question: what happened outside the process? External auditors have been required to ask it ever since — the fraud standards (PCAOB AS 2401 and its AICPA sibling) mandate testing journal entries for evidence of management override precisely because of cases like this. Internal audit inherited the technique and, done well, runs it deeper than the external team ever will: continuously, across the full ledger, with institutional knowledge of who should never be posting what. This guide is the complete internal-audit version: population discipline, the risk-scoring criteria library, selection strategy, what to actually test, and the special problem of top-side entries. It is the sibling of our accounts payable program — AP is how money leaves through the front door; the GL is how the books lie about it.

This guide was rewritten in September 2026 to add what the August version left out: how to size the engagement, Lakeshore Bancorp’s year-end run from population to findings, the disposition grid the examination is recorded in, the findings written in five-Cs form, and the Standards references. The criteria library and the three-layer selection strategy are unchanged. The economics have moved on since the 2024 figures quoted below: the ACFE’s Occupational Fraud 2026 report puts financial statement fraud in 6 percent of its 2,402 cases with a median loss of one million dollars and a median duration of 24 months, and the mechanisms are worked through case by case in the financial statement fraud guide.

In this guide

Why JE testing is the anti-override control

Every process control you test — approval matrices, three-way matches, reconciliations — shares a quiet assumption: transactions flow through the process. Management override is the risk that someone senior enough simply does not use the process. They cannot avoid one thing, though: if they want the financial statements to say something false, a journal entry has to put it there. The GL is the choke point. That is what makes JE testing categorically different from the rest of your program — it is a detective control aimed at the residual risk that survives every other control’s effectiveness, which is why fraud standards treat it as mandatory regardless of how good the control environment looks. The ACFE’s 2024 Report to the Nations supplies the economics: financial statement fraud is only about 5% of occupational fraud cases but carries the highest median loss, $766,000 — and the headline cases run to billions.

The design consequence: JE testing must assume the adversary knows the controls. Sampling randomly from millions of system-generated postings is theater — the interesting entries are, by construction, the ones engineered to look uninteresting or posted where nobody looks. So the method is targeted: define the full population, carve out what cannot be override (pure system postings), score what remains against risk criteria, and put human eyes on the worst-scoring entries plus every entry in the categories where override lives — manual, post-close, and top-side.

Population first: the completeness proof everyone skips

A JE test on an incomplete population is worse than no test — it produces assurance about exactly the entries someone wanted you to see. Before any scoring: prove the extract is the ledger. The proof has three legs. First, roll the extract to the financials: sum debits and credits by account and period and reconcile to the trial balance movement, which must itself tie to the reported statements; if the extract explains every balance movement, nothing posted outside it. Second, inventory the entry sources: subledger interfaces (AP, payroll, billing), automated allocations and reversals, manual entries in the ERP, spreadsheet-uploaded batches, and — the dangerous tail — consolidation-level and post-close adjustments that may live in a different system entirely. Ask explicitly: “show me every way a number can reach the reported financials,” and keep asking until the consolidation system, the reporting tool, and the disclosure process are all in the answer. Third, test the extraction itself: who ran it, from what system, with what parameters — the same information-produced-by-the-entity discipline our model workpaper applies to a termination report applies tenfold to the ledger. WorldCom’s entries were in the ledger; a population pulled “excluding consolidation adjustments” would have missed the fraud by design.

Then segment. Three sub-populations get different treatment: system-generated entries (test the interface and configuration once — sampling identical postings adds nothing); routine manual entries (accruals, reclasses, allocations — the scoring pool); and non-routine entries — post-close, top-side, consolidation, and anything touching the accounts where judgment lives (reserves, impairments, capitalization). The third group is small and gets disproportionate attention, because it is where every famous fraud lived.

Sizing the engagement: hours, data and sequence

A first journal entry engagement is mostly data work, and functions that budget it as a sampling exercise run out of hours before they have examined an entry. The budget below is for a first run at an organisation of Lakeshore’s scale (about 1.9 million ledger lines a year, of which 46,000 lines in 11,200 entries are manual); a company with a tenth of the volume needs perhaps half the hours, because the completeness proof and the scoring build do not shrink with the population. The second year costs a fraction of the first if the pipeline is kept.

PhaseHours (first run)What happensWhat it produces
Planning and data request30Sources inventoried (“every way a number reaches the financials”); fields specified; extraction parameters agreed and witnessed; close calendar and approval thresholds obtainedData request; source inventory; planning memo
Completeness proof40Extract rolled to trial balance movement and to the reported statements; top-side and consolidation entries added from the reporting tool; extraction tested as information produced by the entityReconciliation workpaper; population record
Segmentation and scoring build60System, routine manual and non-routine sub-populations separated; criteria library coded; weights calibrated on one quarter; parameters stored in a configuration tableScoring engine; calibration note
Selection and examination120Layer one (categorical sets) inspected in full; layer two (top scores) examined individually; layer three random sample drawn with a seed; four questions asked of every selected entryDisposition grid; exceptions list
Top-side and consolidation30Population obtained from the consolidation system directly; examined entirely; volume and direction trended; posting access testedTop-side workpaper
Approval control tests40Configuration (separation, self-approval block, delegation rules), access list, and operation for a sampleControl conclusions; access reduction list
Reporting40Findings in five-Cs form; deficiency evaluation where SOX applies; audit committee summary; hand-off of any referral under the protocolReport; issue log entries
Total360Senior auditor 200 hours, analytics auditor 120, manager review 40Quarterly runs thereafter: 60 to 80 hours

Two sequencing rules protect the hours. Do not begin scoring until the completeness proof is signed, because a scoring engine built on an extract that turns out to be missing the consolidation entries has to be rebuilt, and at Lakeshore the first year’s run missed the 140 top-side entries until the reconciliation to the reported statements refused to balance. And examine layer one before layer two, because the categorical sets (top-side, post-close, executive-posted) are where the findings are, and a function that spends its examination hours on the scored tail first may never reach them.

The risk-scoring criteria library

Each criterion earns its place because it correlates with how override actually behaves — frauds need to move numbers late, in quiet corners, through hands senior enough to skip the process. Score entries against the library (one point per hit, weighted where noted), then let the distribution nominate your selections:

#CriterionWhy it signals
1Posted in the close window or after (back-dated to the period)Frauds happen when the target number is known — late in or after the period; WorldCom’s capitalization entries were post-close
2Round numbers at scale ($500,000.00; $3,000,000.00)Real transactions carry cents; invented ones are typed — weight this heavily above a materiality floor
3Just below an approval or review thresholdSame threshold-hugging logic as AP split purchases — people game the lines they know
4Seldom-used account combinationsAccounts touched twice a year draw no reviewer instinct; pairings nobody expects (expense → asset) are the WorldCom signature
5Suspense, clearing, and intercompany accountsThe ledger’s junk drawers — where amounts wait, net, and vanish from scrutiny
6Senior preparer (controller and above), or preparer outside accountingOverride is by definition senior; executives posting their own entries is the single loudest flag
7Missing, generic, or keyword descriptions (“per management,” “plug,” “true-up,” “adjust,” “reclass per…”)Language is evidence; entries with something to hide say little — or say it in tells
8No approver, self-approval, or approver junior to preparerThe approval control’s residue; inverted seniority means the “approval” was administrative
9Weekend, holiday, or odd-hours postingCircumvention prefers empty offices — calibrate to the close calendar before treating as anomalous
10Entries that later reverse (especially straddling a period end)Post it, report it, reverse it — the oldest earnings-management move; pair period-end entries with their reversals
11Credits to revenue or debits reducing reserves late in periodDirectionality matters: entries that improve results at the deadline deserve more skepticism than ones that worsen them
12Spreadsheet-uploaded batchesBulk uploads bypass field-level edit checks and bury single bad lines among hundreds of clean ones
13Entries to entities or ledgers outside the preparer’s remitA US controller posting to the Irish entity is either a process quirk or a place to hide — find out which
14Estimate and judgment accounts: reserves, allowances, impairments, capitalizationWhere GAAP gives discretion, override wears a defensible face — test the support, not just the arithmetic

Calibrate before you trust it: run the scoring on a quarter, examine the distribution, and tune weights so the top tier is dominated by genuinely unusual entries rather than the payroll accrual that innocently trips three criteria every month. A criteria library that flags the same routine entries every period trains everyone — including you — to ignore it.

From scores to selections: the three-layer strategy

Selection is stratified, not random — with one deliberate exception. Layer one: 100% of the categorical high-risk sets. All top-side and consolidation entries, all post-close adjustments above a floor, all entries by executives — these populations are usually small enough to examine completely, and no score should exempt them. Layer two: the top of the risk-score distribution — the 25–50 highest-scoring entries, examined individually. Layer three: a small random sample from everything else (see the sampling guide for sizing, seeded via mySampler) — not because you expect hits, but because a purely criteria-driven selection is predictable, and predictable tests teach a knowledgeable adversary exactly what to avoid. The random layer is the unpredictability premium; say so in the memo, because it is also the honest answer to “why did you pull my boring accrual?”

Testing the selected entries

For each selected entry, four questions — in this order, because the first is the one that catches frauds. Business purpose: can someone explain, in operational terms, why this entry exists? Not “it books the adjustment” (that restates the entry) but what real-world event required the ledger to change. An entry whose purpose cannot be narrated by anyone below the person who ordered it is the profile of every override case ever prosecuted. Support: does documentation exist, is it contemporaneous, and does it actually compute to the amount — a memo asserting a number is assertion, not support. Authorization: approved per policy, by someone independent and senior enough, before or reasonably after posting — and watch for approval theater, the approver who cleared 400 entries in an afternoon. Accounting propriety: right accounts, right period, right treatment — capitalization entries get the extra question WorldCom teaches: does this expense genuinely create a future economic benefit, or does it just create future earnings?

Document each entry’s disposition in a one-line-per-entry grid — purpose confirmed with whom, support reference, conclusion — with the same reperformable discipline as our model workpaper. Where an answer is unsatisfying, resist the urge to average it away: one entry with no articulable business purpose is not a 4% exception rate, it is a lead.

Top-side and consolidation entries: the WorldCom lesson

Top-side entries — adjustments made at the consolidation or reporting level, above the operating ledgers — deserve their own section because they combine maximum power with minimum control. They are posted by the small senior team that runs the close; they frequently live outside the ERP’s approval workflow, in the consolidation tool or a spreadsheet; they can rewrite any number after every process control has finished running; and they are invisible to entity-level reviewers. Every one of those properties is why reserve manipulation and last-mile earnings management happen here. The program: obtain the complete top-side population directly from the consolidation system (not from the team that posts them); expect it to be small — dozens, not thousands — and examine it entirely; demand entry-level support and business purpose exactly as above, with zero deference to seniority; trend the volume and direction quarter over quarter (a rising count of late top-side adjustments that always help results is a finding even when each entry individually survives); and test the access list — who can post at consolidation level is a control test all by itself.

Worked example: Lakeshore Bancorp’s year-end run, from population to findings

Lakeshore Bancorp is the 9-billion-dollar regional bank used across this site, with a 22-person internal audit function, SOX materiality of 6.5 million dollars, a CFO review threshold for manual entries of 250,000 dollars and a second approver required above 50,000. The function runs the scoring engine from the journal entry analytics catalog at each quarter-end as part of the SOX program’s fraud and override coverage, with the year-end run the largest. The table follows one year-end run through the method above, from the population to the findings, and the outcome of each step is the point.

StepWhat happened at LakeshoreWhat it shows
PopulationAbout 1.9 million ledger lines; 46,000 lines in 11,200 manual entries; 140 top-side and consolidation entries in the reporting tool. The extract reconciled to trial balance movement exactly and to the reported statements only after the 140 top-side entries were added; the first year’s run had missed them, which was that run’s own first finding.The completeness proof finds the hole in the population before the population finds it for you
Scoring: above 15 (38 entries, 61 million dollars)Inspected in full: 31 appropriate and supported; 5 appropriate but unsupported; 2 inappropriate, a 1.1-million-dollar reclassification between fee income lines posted after the hard close with a generic description, and a 640,000-dollar manual true-up to the allowance posted outside the estimation process; both corrected.The top band is small enough to examine entirely and is where the accounting errors were
Scoring: 10 to 15 (212 entries, 148 million dollars)Sample of 40 stratified by the criterion that drove the score: 36 appropriate and supported, 4 unsupported, no errors.The middle band produces documentation findings rather than accounting ones
Scoring: 5 to 10 (1,460 entries, 390 million dollars)Sample of 25 plus pattern analysis: all appropriate; 61 percent of the band’s entries came from three posters in the loan accounting team, a workload concentration referred to the controller.Patterns in the low bands are operational findings, not fraud findings, and still worth reporting
Scoring: below 5 (9,490 entries, 2.1 billion dollars)Random sample of 15 plus pattern analysis: all appropriate; 380 entries carried descriptions consisting only of a person’s initials.The random layer is the unpredictability premium; the pattern analysis pays for it
Mandatory inspections140 top-side, 26 by senior finance management, 9 by workflow administrators, 84 million dollars, inspected in full regardless of score: all top-side entries supported; two entries by the Chief Accounting Officer to accrual accounts approved by a direct report.Categorical sets are examined whole because score does not exempt seniority
The five-point tests14 self-approved entries, all through a delegation the workflow allowed when the approver was on leave; six entries by users after their termination dates, through a shared close-team login; 88 post-close entries, 71 authorised late adjustments on the close checklist and 17 not, including the fee income reclassification.The approval control’s residue is where configuration findings live
Where it wentThe delegation configuration and the shared login evaluated as deficiencies in the operation of the approval control; the post-close channel as a design gap in the close process; the Chief Accounting Officer’s entries as an entity-level segregation matter for the audit committee; all under the method in the deficiency evaluation guide.A journal entry run ends in a deficiency evaluation, not in a list of anomalies

Three things about Lakeshore’s run generalise. No entry in the run was fraudulent, and the run was still the most valuable engagement in the SOX program that year, because it found a 1.1-million-dollar accounting error, an allowance adjusted outside its process, a delegation rule that turned self-approval on whenever an approver went on holiday, a shared login used by people who no longer worked there, and a segregation issue at the top of the finance function; every one of those is a condition under which a fraud would have looked exactly like an ordinary entry. The categorical inspections found the segregation matter and the five-point tests found the configuration findings; the score found the accounting errors; the three layers found different things, which is the argument for all three. And the run’s first finding, in its first year, was about itself: the population was incomplete, and the reconciliation to the reported statements was the only thing that said so.

Testing the approval control itself

Entry-level testing tells you what happened; control testing tells you what can happen. Three checks complete the picture. Configuration: does the system enforce preparer-approver separation, block self-approval (including through delegation), and require approval before posting — or is approval a checkbox after the fact? Evidence the settings, then test change control over them. Access: who holds manual-JE posting rights, and does the list match the roles that need it? Every JE audit should shrink this list; posting access accretes like sediment. Operation: for a sample of entries (layer three doubles here), was the approval genuine — the right person, engaging with support, within a plausible time? Segregation failures found here compound everything else: a preparer-approver conflict plus spreadsheet-upload access is the full override toolkit in one person’s hands.

The disposition grid: one line per entry

The examination is recorded in a grid with one line per selected entry, and the grid is the workpaper a reviewer, an external auditor relying on the work, or an investigator will read. The template below is the column set; it lives in the analytics workbook with the population record and the scoring parameters, and its totals tie to the selection memo. Keep the conclusions in the fixed vocabulary, because “supported” and “appropriate” are different questions and a grid that blends them cannot be aggregated.

Journal entry disposition grid — run [period]; population record [reference]; scoring parameters [version]; selection memo [reference]; prepared by [name, date]; reviewed by [name, date].

Columns. 1. Entry ID and line. 2. Selection layer (categorical set named; score band; random). 3. Score and the criteria that drove it. 4. Poster, approver, approval timestamp, delegation flag. 5. Accounts and amount. 6. Business purpose as explained, and by whom (name and role); “restates the entry” recorded as no explanation. 7. Support reference, date of support relative to posting, and whether it computes to the amount. 8. Authorisation: per policy, independent, senior enough, timing. 9. Accounting propriety: accounts, period, treatment; for capitalisation, the future-benefit question. 10. Conclusion, one of: appropriate and supported; appropriate, unsupported; inappropriate accounting; unable to conclude. 11. Follow-up: correction posted, control finding reference, referral under the protocol, or none. 12. Reviewer note.

Totals. By layer and by conclusion; value of entries by conclusion; exceptions carried to the findings list with their references. An entry with no articulable business purpose is listed under follow-up as a lead, never averaged into a rate.

Coordinating with the external auditor without merging with them

The external auditor runs its own journal entry procedures under AS 2401 or ISA 240 every year-end, and the two programs should be coordinated without becoming one. Share the completeness reconciliation, because two teams proving the same extract complete is wasted effort and the external auditor may accept the function’s proof as evidence under its reliance standards; share the criteria library and the calibration note, because the external team’s criteria are usually a subset and the comparison is instructive both ways; and agree the timing so that the function’s run precedes the year-end audit and the external team can direct its own examination at what the function did not cover. Do not share the selections in advance, because a selection known to management before the entries are examined is a selection management can prepare for; do not let the external auditor’s coverage substitute for the function’s own, because the external team examines for the financial statements at materiality and the function examines for control and fraud at the level of a single entry; and do not run the function’s program only at year-end because the external auditor does, because the override that matters happens in the quarters nobody audits. Where the external auditor wants to rely on the function’s work, the file has to meet their standard, which is the reason the disposition grid records who explained each entry’s purpose and where the support sits rather than a tick.

Building it as an analytic — once, then continuously

Everything above is expressible as data logic over a handful of fields: entry ID, line, account, amount, debit/credit, posting and effective dates, source, preparer, approver, description, entity. Which means the second year costs a fraction of the first: keep the completeness reconciliation, the scoring engine, and the disposition grid as a repeatable pipeline, and the annual JE audit becomes a quarterly — then monthly — monitoring routine, with human review reserved for the scored tail. That is the natural on-ramp to continuous auditing (our continuous auditing versus continuous monitoring guide maps the operating models), and it changes the deterrence math: a scoring engine that runs monthly is a control everyone knows is watching, which is worth more than the findings it produces. One warning as you automate: keep the completeness proof in the pipeline permanently. An automated JE monitor fed by an incomplete extract is the most convincing false assurance a function can build.

The findings that recur, and wording that lands

Journal entry findings are read by the people who post and approve the entries, so they land only in five-Cs form with a count and a criterion. Delegation configuration (“14 manual entries in the quarter were approved by their own posters through the workflow’s leave-delegation rule, which transfers approval rights to the poster’s delegate without excluding the poster; the manual journal policy prohibits self-approval in any form”). Shared credentials (“six entries were posted after their nominal users’ termination dates through a shared close-team login; the access policy requires named credentials for posting rights”). Post-close channel (“88 entries were posted after the hard close, 17 of them not on the authorised late-adjustment list, including a 1.1-million-dollar reclassification of fee income with a generic description; the close procedure permits post-close entries only from the checklist with controller approval”). Estimates outside process (“a 640,000-dollar manual adjustment to the allowance was posted outside the estimation process and its governance; the allowance methodology requires all adjustments to flow through the quarterly calculation and its review”). Segregation at the top (“two entries by the Chief Accounting Officer to accrual accounts were approved by a direct report; the policy requires approval by someone senior to the poster”). And descriptions (“380 entries carried descriptions consisting only of initials; the policy requires a description sufficient for an independent reviewer”). Write the cause as the configuration, the procedure or the reporting line that produced the condition, and follow the root cause guide; where the cause is a person’s intent, the matter leaves the report for the protocol. The Global Internal Audit Standards frame the engagement’s obligations: Standard 13.2’s consideration of fraud in planning is what puts the override lens on the plan, and Standard 14.1’s requirement for sufficient, reliable, relevant information is what the completeness proof satisfies.

Where to go next

Journal entry testing is internal audit’s answer to the question every other test quietly avoids: what if the people above the controls are the problem? Run it with the full discipline, a proven-complete population, criteria that mirror how override actually behaves, stratified selection with a random layer, business-purpose questioning that defers to no one, and total coverage of the top-side tier, and you hold the choke point WorldCom proved matters: whatever the scheme, the entries have to exist. The scoring engine and its twenty-five tests are in the journal entry analytics catalog; the mechanisms the entries serve are in the financial statement fraud guide; what to do when an entry has no innocent explanation is in the first-48-hours protocol; the deficiency logic for what the run finds is in the control deficiency evaluation guide; and the rest of the disbursement-fraud universe starts with how to audit accounts payable.

Related guides

Comments

Leave a Reply

Discover more from internalauditguide.com

Subscribe now to keep reading and get access to the full archive.

Continue reading