Every financial statement fraud that has ever been unwound ran through the general ledger as journal entries, and most of the entries looked, individually, like every other entry. That is why journal entry testing exists as a standalone requirement in the auditing standards, and why the interesting work is not in inspecting a sample of entries but in deciding which entries deserve inspection. This catalog sets out twenty-five tests in four families, timing, amounts, accounts, and posters and approvals, each with its logic, the data it needs, a starting threshold, what a hit usually means and what produces false positives, then combines them into a composite risk score that ranks the population, and shows how the score drives the procedure that follows. It includes the data preparation, a scoring and triage record template, a worked scoring run at a bank, and the path from an annual exercise to a monthly one.
This guide was rewritten in September 2026 from a shorter version published on 1 September 2026. The external auditor’s obligation sits in PCAOB AS 2401, whose paragraphs on journal entries and other adjustments require procedures to test the appropriateness of entries recorded in the general ledger and adjustments made in preparing the financial statements, and list the characteristics of entries that warrant attention: entries to unrelated, unusual or seldom-used accounts; entries by individuals who typically do not make entries; entries recorded at period-end or as post-closing entries with little or no explanation; entries made before or during preparation of the statements that lack account numbers; and entries containing round numbers or consistent ending numbers. Each of those characteristics is a test below, and the composite score is the way of applying all of them at once.
In this guide
- The data you need first, and the completeness proof
- Family 1: Timing (7 tests)
- Family 2: Amounts (6 tests)
- Family 3: Accounts (6 tests)
- Family 4: Posters and approvals (6 tests)
- Building the baseline
- The composite risk-score model
- Triage: from score to procedure
- Management override and the entries the score will not find
- The scoring and triage record
- Worked example: Lakeshore Bancorp’s year-end scoring run
- Coordinating with the external auditor
- From annual exercise to monthly program
- Common mistakes
- Related guides
The data you need first, and the completeness proof
The population is every journal entry line posted to the general ledger in the period, from every source, plus the consolidation and top-side adjustments made outside the ledger in the reporting tool or in spreadsheets, which are the entries most often missing from a “complete” extract. The fields: entry ID, line number, company code, account, cost center or profit center, amount and sign, currency, posting date, effective or document date, entry date and time, source or document type (manual, automated from a sub-ledger, recurring, reversal, allocation), poster, approver and approval timestamp, description, reference, reversal indicator and reversed entry ID. Supporting tables: the chart of accounts with account type and status, the user master with role and department, the HR file for termination dates, the approval workflow configuration and the period close calendar.
Completeness is proved three ways before any test runs: the sum of all lines by account reconciles to the trial balance movement for the period; the sum by period reconciles to the ledger’s own posting totals; and the top-side and consolidation entries reconcile to the difference between the ledger trial balance and the reported financial statements, which is the reconciliation that finds the entries nobody extracted. The population record, with the queries and the reconciliations, goes in the workpaper following the IPE testing guide. Manual entries are then separated from system-generated ones by source type, because most tests run on the manual population and use the automated population as the baseline, and the source-type mapping itself is verified against the ledger’s configuration rather than assumed.
Two data problems recur at extraction and are worth solving before the first run rather than during it. Multi-ledger environments (a bank with a separate ledger for a mortgage subsidiary, a group with entities on different systems) need the population assembled across ledgers with a common account mapping, and the intercompany mirror test cannot run at all until that mapping exists. And ledgers that store the approver in a workflow system rather than on the entry need the workflow log joined to the entries by document number, with the join rate reported: an entry population where 8 percent of manual entries cannot be matched to a workflow record has an 8 percent hole in every approval test, and the hole is the first finding.
Family 1: Timing (7 tests)
Timing tests operationalize the standard’s “recorded at the end of the period or as post-closing entries” characteristic, and they need the close calendar as an input: the last day of the period, the soft-close and hard-close dates, and the date the financial statements were finalized, for each period.
| Test | Logic | Data and threshold | A hit usually means | False positives |
|---|---|---|---|---|
| T1 Post-close entries | Manual entries with an effective date in a period, posted after that period’s hard close, or after the statements were finalized | Posting date vs close calendar | Adjustments made after review, the classic override channel; or a close that is not actually closed | Authorized post-close adjustments with documented approval (which should be few and listed) |
| T2 Period-end clustering | Manual entry count and value in the last two business days of the period and the first five after, relative to the period’s daily average | Posting dates; daily volumes | Close pressure; entries made to hit a number; or normal close activity, which the baseline separates | A close process concentrated by design; compare to prior periods |
| T3 Weekend and holiday posting | Manual entries posted on non-working days | Posting date vs calendar | Activity outside supervision, or a batch job with a system user | Global teams in other calendars; scheduled processes |
| T4 Off-hours posting | Manual entries posted outside working hours by human users | Entry timestamps; user type | Same as T3 at finer grain; a poster working alone | Time zones; close-week overtime |
| T5 Entry-date versus effective-date gaps | Difference between posting date and effective date beyond a threshold (for example 30 days), especially backdated | Both date fields | Backdating to move results between periods | Late invoices legitimately accrued back |
| T6 Reversal timing | Entries reversed in the following period within days of the close, or reversed and re-posted with changed amounts; reversals of reversals | Reversal indicators and dates | Temporary entries to dress a period-end balance | Standard auto-reversing accruals, excluded by source type |
| T7 Final-hours metric movers | Entries in the last hours before the close that move a reported metric (margin, covenant ratio, segment result) across a threshold | Entry timing; metric definitions and thresholds | Entries made with the metric in mind | Genuine late adjustments; the test flags for inspection, not conclusion |
Family 2: Amounts (6 tests)
| Test | Logic | Data and threshold | A hit usually means | False positives |
|---|---|---|---|---|
| A1 Round amounts | Manual entry lines at round thousands or hundreds of thousands, as a share of a poster’s or an account’s lines against the population | Line amounts | Estimated or invented amounts; the standard’s “round numbers” characteristic | Accruals legitimately estimated in round amounts; allocations |
| A2 Just-under-threshold | Entries in the band just below approval or review thresholds (a second approver above $50,000, a CFO review above $250,000) | Line and entry amounts; approval thresholds | Structuring to avoid review | Natural amounts near round thresholds |
| A3 Benford deviation | First-digit and first-two-digit distribution of manual entry amounts by poster and by account against Benford, for strata with enough lines | Amounts; strata of several hundred lines or more | Invented amounts; a poster whose numbers do not arise naturally | Small strata; accounts with fixed amounts |
| A4 Recurring identical amounts | The same amount posted repeatedly by one poster to one account, outside recurring-entry source types | Amounts by poster and account | Manual repetition of what should be a recurring entry; or a fixed fictitious charge | Legitimate fixed monthly charges posted manually |
| A5 Offset asymmetry | Entries whose debit and credit sides land in unrelated areas (an expense against an unrelated balance sheet account), or one-sided entries balanced through suspense | Line pairs within entries; account mapping | Concealment through an unexpected offset; the “consistent ending numbers” and “unrelated accounts” characteristics together | Legitimate reclassifications; complex allocations |
| A6 Split entries | Several entries by one poster within a short window, same description or accounts, whose total exceeds a threshold each avoids | Entry sequences; thresholds | Structuring | Batch postings by design (a sub-ledger interface split by document) |
Family 3: Accounts (6 tests)
Account tests operationalize “unrelated, unusual or seldom-used accounts”, and they depend on a baseline of what is normal for each account: its usual source types, its usual posters, its usual counter-accounts, built from the automated population and the prior year’s manual entries.
| Test | Logic | Data and threshold | A hit usually means | False positives |
|---|---|---|---|---|
| C1 Seldom-used accounts | Manual entries to accounts with fewer than N postings in the prior twelve months, or with no manual postings historically | Account activity history | An account chosen because nobody looks at it | New accounts opened for a new product or entity |
| C2 Suspense and clearing activity | Manual entries to suspense, clearing and unapplied accounts; aged balances in them; entries that move balances between suspense accounts | Account type mapping; balances by age | Parking of differences; concealment in transit accounts | Normal clearing that reverses within days; the age filter separates it |
| C3 Intercompany without a mirror | Intercompany entries with no matching entry in the counterparty entity in the period, or with mismatched amounts | Intercompany accounts across entities | Profit moved between entities; eliminations manipulated; or an out-of-balance intercompany process | Timing differences resolved at consolidation |
| C4 Reserve and estimate accounts hit manually | Manual entries to allowance, reserve, accrual and valuation accounts outside the documented estimation process, or by posters outside the estimation team | Account mapping to estimates; poster roles | Cookie-jar adjustments; the highest-risk hit in the family | Approved true-ups from the estimation process, which should be traceable to the estimate file |
| C5 Closed-account resurrection | Entries to accounts flagged closed, blocked or inactive in the chart | Chart of accounts status | Account status controls not enforced; an old account reused for concealment | Reopening authorized for a specific correction |
| C6 Unusual account pairings | Debit-credit account combinations never seen in the automated population or the prior year, weighted by amount | Line pairs; baseline of historical pairings | An entry doing something the process never does | Genuine new transaction types; the baseline needs annual refresh |
Family 4: Posters and approvals (6 tests)
| Test | Logic | Data and threshold | A hit usually means | False positives |
|---|---|---|---|---|
| U1 Unexpected posters | Manual entries by users outside the accounting roles, including executives, IT users and service accounts | Poster vs role and department | The standard’s “individuals who typically do not make entries”; the management override pattern | Legitimate posting by system administrators during migrations, which is itself worth a finding |
| U2 Self-approved entries | Approver equals poster, or approval by a user whose approval right was granted by the poster | Workflow log; role administration log | Segregation failure; routing or delegation configured to allow it | None; every hit is at least a control finding |
| U3 SoD-clash posters | Posters who also hold conflicting rights: vendor master, cash disbursement release, bank reconciliation, chart of accounts maintenance | Poster list vs access rights | A user who can create and conceal | Small finance teams where the conflict is known and compensated, which is the finding |
| U4 Departed-user postings | Entries by users after their HR termination date, or by generic and shared IDs | Poster vs HR terminations; user master | Access not removed; shared credentials in use | Termination dates recorded late |
| U5 Poster-volume anomalies | Posters whose manual entry count or value in the period is far above their own history or their peers’ | Poster history | Someone doing unusual work; a workload that hides an entry | Role changes; project work |
| U6 Description quality | Entries with blank, generic (“adjustment”, “correction”, “per JS”) or duplicated descriptions, or descriptions that do not match the accounts hit | Description text; simple text rules | The “little or no explanation” characteristic; entries nobody wanted to explain | Cultural brevity; the test ranks, it does not conclude |
Building the baseline: what normal looks like for each account and poster
Half the catalog compares an entry to what is normal, and normal has to be computed before it can be compared against. The baseline is built from two populations: the automated entries of the current period, which show what the processes do by design, and the manual entries of the prior twelve months, which show what the people do by habit. For each account it records the source types that post to it, the posters who post to it manually and how often, the counter-accounts it is paired with and the amount distribution of its manual lines. For each poster it records the accounts they touch, their monthly entry count and value, their usual posting hours and their approval relationships. The seldom-used account test, the unusual pairing test, the unexpected poster test and the poster-volume test all read from this table, and their false-positive rates depend entirely on its quality: a baseline built from a year in which the organization changed its chart of accounts or reorganized its finance team will flag the new normal as abnormal for a year.
Three rules keep the baseline honest. It is refreshed annually, after the year-end run, so that the current year’s manual entries become next year’s norm only after they have been tested; a baseline that absorbs untested entries teaches the model that last year’s improper entry is normal. Accounts and posters new in the period get no baseline and are flagged as such, rather than being treated as anomalous or as normal; the new product’s accounts and the new controller’s entries are inspected in full in their first period and enter the baseline afterward. And the baseline excludes entries that were corrected, reversed as errors or referred, so that the model does not learn from the population it is meant to catch. The baseline table is itself information produced by the entity and gets a completeness check against the ledger, as the IPE testing guide would require of any other input.
The composite risk-score model
No single test finds a well-constructed improper entry, because a careful poster avoids any one characteristic. A composite score finds the entry that has several. The model is deliberately simple, so that it can be explained to an audit committee and re-performed by an external auditor: each test contributes points when it fires, the points are weighted by how diagnostic the test is, and an entry’s score is the sum across the tests it hits, with a multiplier for amount so that a high-scoring $200 entry does not outrank a moderately scoring $2 million one. The weights below are a starting point that the program calibrates from its own results: after each run, the tests whose hits were mostly false positives lose weight and the tests that contributed to corroborated findings gain it.
| Component | Points | Rationale |
|---|---|---|
| U2 self-approved; U4 departed user; C5 closed account; T1 post-close | 5 each | Control failures by definition; each alone justifies inspection |
| C4 reserve or estimate hit manually; U1 unexpected poster; C3 intercompany without mirror; T7 metric mover | 4 each | The management override channels; high diagnostic value |
| A5 offset asymmetry; C6 unusual pairing; C1 seldom-used account; U3 SoD-clash poster; T6 reversal timing | 3 each | Concealment patterns; moderate false-positive rates |
| A2 just-under-threshold; A6 split; T5 date gap; C2 suspense; U6 poor description | 2 each | Common in improper entries and in careless ones |
| A1 round; A3 Benford; A4 recurring identical; T2 clustering; T3 and T4 timing; U5 volume | 1 each | Weak individually; useful in combination |
| Amount multiplier | 1.0 below the review threshold; 1.5 above it; 2.0 above performance materiality; 3.0 above overall materiality | Keeps large entries at the top; based on the entry’s absolute value, not its net effect |
Scores are computed per entry (the header level), with line-level tests rolled up so that an entry with one suspicious line carries the line’s points. The output is the manual population ranked by score, with the distribution reported: how many entries scored above 15, above 10, above 5, and the value in each band. The distribution is itself a finding when it moves year over year, and it is the number the audit committee can follow. The journal entry testing guide covers the procedures applied to the entries selected; this catalog is the selection engine.
Triage: from score to procedure
The score determines the procedure, and the program writes the mapping down before the run so that the selection is not adjusted to fit the budget afterward. Entries above the top threshold are inspected in full: the supporting documentation, the business rationale from someone other than the poster, the approval, the accounting treatment, and where the entry affects an estimate, the estimate file. Entries in the middle band are inspected on a sample sized to the band’s population, with the sample stratified by the test that drove the score so that each family is represented; the sampling memo template records the design. Entries below the bottom threshold are not inspected individually but their distribution is analyzed for patterns, by poster, account and period, and the patterns become findings about the process rather than about entries. Every hit on the five-point tests is inspected regardless of score, because a self-approved entry is a control failure whatever else it is.
Corroboration follows the entry’s purpose. An entry explained as a reclassification is corroborated by the two accounts’ subsequent behavior; an accrual by the invoice or the calculation that followed; an estimate adjustment by the estimate file and the approver’s independent review; an intercompany entry by the counterparty’s books. The conclusion for each inspected entry is one of: appropriate and supported, appropriate but unsupported (a documentation finding), inappropriate accounting (an error), or improper (referred under the fraud response protocol described in the fraud risk management guide), and the counts by conclusion are the summary the report leads with.
Management override and the entries the score will not find
The score finds entries with characteristics. It will not find an entry posted by the right person, at the right time, to a usual account, for a plausible amount, with a good description, that is nonetheless wrong, and that is what a competent override looks like. Three procedures cover the gap; the first is the “other adjustments” AS 2401 explicitly brings into scope, and the second follows from its attention to entries by individuals who do not typically make them. Top-side and consolidation entries are inspected in full every period, because they bypass the ledger’s controls by construction and their population is small. Entries by senior finance management and by anyone with the ability to change the workflow are inspected in full, whatever their score, because the score’s poster tests treat them as expected posters. And a sample of low-scoring entries is inspected every period, small and random, so that the scoring model is tested against the population it dismisses; if the random sample turns up an inappropriate entry, the model has a gap and the weights are revisited. The SOX scoping guide sets out where the override risk sits in the program’s design, and the fraud red flags guide the non-data indicators that should change the weights for a period, such as a missed forecast, a covenant close to breach or a departing CFO.
The scoring and triage record
Journal entry analytics run record
1. Run. Period(s) covered; run date; analyst; ledger(s) and entities in scope; reporting tool and top-side sources included.
2. Population. Line and entry counts by source type; reconciliation to trial balance movement, to ledger posting totals and to the reported statements (top-side reconciliation); source-type mapping verified against configuration; close calendar used.
3. Parameters. Thresholds for each test; approval and review thresholds with source; materiality and performance materiality; baseline period for account and poster norms; test weights and amount multipliers in force.
4. Results. Hits per test with counts and value; score distribution by band with counts and value; entries inspected in full (list); sample design for the middle band; random low-score sample; mandatory inspections (top-side, senior management, workflow administrators).
5. Inspections. Per entry: score and contributing tests; support obtained; rationale source; approval; conclusion (appropriate and supported / appropriate but unsupported / inappropriate accounting / improper); amount; correction or referral.
6. Findings. Control failures (self-approval, departed users, closed accounts, post-close channel) with counts; process findings from the low-score pattern analysis; estimate and intercompany findings; referrals.
7. Calibration. Tests whose hits were mostly false positives and the weight change; tests that contributed to findings; random-sample result and model conclusion; parameters for the next run.
8. Sign-off. Analyst, reviewer, dates; coordination note with the external auditor where the population or the scoring was shared.
Worked example: Lakeshore Bancorp’s year-end scoring run
Lakeshore Bancorp, the illustrative $9 billion regional bank used across this site, posts about 1.9 million ledger lines a year, of which roughly 46,000 lines in 11,200 entries are manual, plus 140 top-side and consolidation entries in the reporting tool. Internal audit runs the catalog at each quarter-end as part of the SOX program’s fraud and override coverage, with the year-end run the largest. The population reconciled to trial balance movement exactly and to the reported statements after the 140 top-side entries were added; the first year’s run had missed them, which was the run’s own first finding. Materiality was $6.5 million and the CFO review threshold for manual entries $250,000, with a second approver required above $50,000.
| Score band | Entries | Value (absolute) | Procedure | Result |
|---|---|---|---|---|
| Above 15 | 38 | $61 million | Inspected in full | 31 appropriate and supported; 5 appropriate but unsupported; 2 inappropriate accounting (a $1.1 million reclassification between fee income lines posted after the hard close with a generic description, and a $640,000 manual true-up to the allowance outside the estimation process, both corrected) |
| 10 to 15 | 212 | $148 million | Sample of 40, stratified by driving test | 36 appropriate and supported; 4 unsupported; no errors |
| 5 to 10 | 1,460 | $390 million | Sample of 25 plus pattern analysis | All appropriate; pattern: 61 percent of the band’s entries were by three posters in the loan accounting team, a workload concentration referred to the controller |
| Below 5 | 9,490 | $2.1 billion | Random sample of 15; pattern analysis | All appropriate; pattern analysis found 380 entries with descriptions consisting only of a person’s initials |
| Mandatory inspections | 140 top-side; 26 by senior finance management; 9 by workflow administrators | $84 million | Inspected in full regardless of score | All top-side entries supported; two entries by the Chief Accounting Officer to accrual accounts approved by a direct report, a segregation finding at the top of the house |
The five-point tests produced their own list: 14 self-approved entries (U2), all through a delegation the workflow allowed when the approver was on leave, which was the configuration finding; six entries by users after their termination dates (U4), through a shared close-team login; no closed-account entries; and 88 post-close entries (T1), of which 71 were the authorized late adjustments listed in the close checklist and 17 were not, including the fee income reclassification. The run’s findings went into the year-end deficiency evaluation described in the deficiency evaluation guide: the delegation configuration and the shared login as deficiencies in operation of the approval control, the post-close channel as a design gap in the close process, and the Chief Accounting Officer’s entries as an entity-level segregation matter for the audit committee. The calibration note dropped the Benford weight, which had produced 900 hits and nothing corroborated, and raised the weight of the post-close test after it drove both errors found.
Coordinating with the external auditor’s journal entry testing
The external auditor will test journal entries whether or not internal audit does, and the two exercises can either duplicate each other or compound. The efficient arrangement has three parts. The population is extracted once, by the organization, with the reconciliation to trial balance movement and to the reported statements performed and documented, and both auditors work from the same reconciled file; that alone removes the most common year-end argument, which is whether the ledger the auditor received was the one the statements came from. The scoring model is disclosed to the external auditor, weights and thresholds included, so that the auditor can decide how much to rely on the selection and where to add their own criteria; auditors typically run their own selection in addition rather than instead, but a disclosed model changes the size of the addition. And the results of internal audit’s inspections are made available with the support obtained, so that an entry inspected and corroborated by internal audit is re-performed by the external auditor on a sample rather than in full. What internal audit does not do is adjust its selection to what the external auditor asks for; the two selections serve different objectives, the SOX program’s override coverage and the financial statement audit’s fraud procedures, and the internal audit file shows a selection made on its own criteria, with the external auditor’s requests recorded separately. The SOX scoping guide describes the same principle for the scope as a whole.
From annual exercise to monthly program
The catalog is coded once, with parameters in a configuration table and the baseline of account and poster norms refreshed annually, and then run monthly on the prior month’s manual population, taking two to four hours after the first year. The five-point tests become management controls: the controller’s team receives the self-approval, departed-user, closed-account and post-close hits before the month is finalized, acts on them, and evidences the action, and internal audit tests that control quarterly instead of performing it, along the line the continuous auditing versus continuous monitoring guide draws. The score distribution becomes a monthly metric in the CFO’s close pack, which is the single change that most reduces high-scoring entries, because posters learn that the description field is read. And the external auditor is offered the population reconciliation and the scoring at year-end, which most accept as the basis for their own selection with their own additions, saving both sides the duplicate extraction. The accounts payable analytics catalog and the payroll analytics catalog follow the same structure.
Common mistakes
Extracting the ledger and forgetting the top-side entries. Running the tests on all entries instead of separating manual from automated, so that the automated volume swamps every distribution. Using thresholds nobody wrote down, then adjusting them to fit the hours available. Scoring without an amount multiplier and inspecting forty small entries. Treating a high score as a conclusion rather than a selection. Inspecting the poster’s own explanation and calling it corroboration. Excluding senior management’s entries because they are “expected posters”. Skipping the random low-score sample, so the model is never tested. Keeping the Benford test at full weight because it is famous. And not sharing the population and the model with the external auditor, so that the same ledger is extracted twice and reconciled once. The journal entry testing guide covers the inspection procedures and the SOX 404 guide the program in which the override coverage sits.
Related guides
- Journal entry testing
- The accounts payable analytics catalog
- The payroll analytics catalog
- Evaluating control deficiencies
- SOX scoping and risk assessment
- SOX 404 explained
- IPE testing
- Audit sampling memo template
- Fraud risk management
- Fraud red flags
- Continuous auditing versus continuous monitoring
- Segregation of duties beyond the ERP
- Financial statement fraud — the mechanisms behind the override entries the scoring criteria are built to find, with Pennine’s quarterly program
- When internal audit finds fraud: the first 48 hours — what to do when a scored entry turns out to be one
Leave a Reply