In most low- and middle-income countries, the binding fiscal constraint is administration, not legislation. The tax laws are on the books. Much of the revenue they mandate doesn’t arrive.
The IMF’s latest assessment of developing-country tax capacity finds that low-income countries are running about 9 percentage points of GDP below what their economies could sustainably yield. Recovered, that gap would lift the group’s average tax-to-GDP from around 13% to roughly 22% — the level emerging markets already manage. Across the $40 trillion in combined GDP of low- and middle-income countries, this amounts to $2–3.5 trillion in untapped revenue each year.
For scale: the OECD’s preliminary numbers put total global official development assistance at $174.3 billion in 2025, its largest single-year fall on record. The tax owed but uncollected across LMICs is at least an order of magnitude bigger than the entire aid pool, and the two have just diverged sharply.
Where the political will to collect exists, the limit is the enforcement capacity that keeps pace with the modern economy.
Some of that capacity has just become technologically tractable. The administrative data tax authorities already hold has crossed thresholds of volume, structure, and interoperability that earlier reformers didn’t have to work with. The bottleneck in collection lies in moving from access to information to the engineering work of using it.
I’m an economist working on tax administration in South Asia, mostly with Pakistan’s federal and provincial revenue authorities. Pakistan anchors the argument here because the data substrate and policy context are unusually visible right now.
What follows traces the gap to its cause and why that cause has only now become fixable. It then lays out what a system to close it looks like in practice, and why such a system pays off only when it is run as an evaluated instrument rather than bought as a product.
1. The arithmetic
Pakistan’s 2023 Economic Census counted roughly 7.2 million economic units operating in the country. The bulk of those (every retailer, service firm, manufacturer, and professional practice) fall under the Inland Revenue: income tax, sales tax, and federal excise tax. FBR’s total sanctioned establishment runs to about 28,000 posts across all cadres and support staff. Within that, the Inland Revenue Service officer cadre charged with assessment and audit (Assistant Commissioner and up) numbers about 1,200 sanctioned posts. Add the field-level Inspectors-IR (BS-16), who carry out site visits, surveillance, and enforcement operations, and the combined IR audit-and-field-enforcement cadre comes to roughly 2,250 sanctioned positions. One sanctioned IR audit-and-enforcement post for every 3,200 economic units.
That ratio is the universal LMIC pattern: a tax universe orders of magnitude larger than the workforce charged with assessing it. India’s GST authority, Bangladesh’s NBR, and Kenya’s KRA all face it. No plausible civil-service expansion closes the gap by hiring. A tenfold expansion of Pakistan’s IR audit and enforcement workforce would still leave each staffer responsible for hundreds of firms. Best, Shah, and Waseem (2021), studying FBR’s randomized-audit program directly, document that even within the much narrower VAT-registered base, the agency could not audit more than 5% of firms a year. That’s the registered subset, well short of the 7.2-million-firm universe.
At around 11% of GDP, federal plus provincial, Pakistan’s total tax take sits in the bottom third of its peer set, below India, Vietnam, Sri Lanka, the Philippines, Kenya, and Rwanda. The World Bank estimates Pakistan’s tax capacity, what its economy could sustain if administration caught up, at roughly 22% of GDP. The gap is on the order of ten to twelve percentage points; in dollar terms, north of US$25 billion a year.
The response everywhere is the same: lean on third parties (banks, large corporations, importers) to remit on behalf of the broader base. In Pakistan, withholding accounts for around 58% of FBR’s income-tax collection, and the advance-tax regimes account for roughly one-third. The two together account for the overwhelming majority. Collection on Demand, the part that follows an actual audit, was about 5% in FY 2024–25, roughly Rs 267 billion, even after more than doubling year-on-year. Voluntary payments alongside returns make up the small remainder.
That adaptation has a ceiling. The alternative is to make more per-taxpayer decisions than any human workforce can process, and for every low-capacity administration, that option was never on the table. Every LMIC made the first choice because the second didn’t exist.
The argument here is that the second option is becoming feasible. The constraint on per-taxpayer enforcement was the cost of integrating, storing, and reasoning over administrative data at the scale of a national tax base. That cost has fallen sharply, and the data itself is now more structured than it was. What’s left is not a data-collection problem but the question of whether the authority builds the layer that turns those flows into enforcement.
2. The data has moved
The standard objection to using AI for tax enforcement in poor countries is that the data is too sparse, too messy, too corrupt. That was a reasonable objection a few years ago. It’s wrong now.
The most consequential development inside Pakistan’s FBR over the last few years has been a quiet data-infrastructure build, less visible than the higher-profile reforms but more substantive. The agency has assembled what amounts to a third-party sensor network across the economy.
Digital invoicing has been rolled out in phases, and more than 45,000 firms now issue invoices through the FBR’s central platform, capturing the bulk of formal-sector B2B transactions. Point-of-sale integration covers around 11,400 Tier-1 retailers, whose roughly 24,000 outlets, together with integrated restaurants and textile/leather chains, bring the connected POS footprint to about 36,000 outlets. AI-based video metering is used on cement, sugar, beverage, and tobacco production lines. Beneath the new build sits a much older and larger store of administrative data that the authority has held for years: sales-tax returns across five revenue authorities, withholding statements covering trillions of rupees in tax flows, customs declarations, records of property transfers, vehicle ownership, and foreign travel. None of these streams is complete: coverage is partial, compliance is uneven, and the pieces still sit largely unlinked across agencies.
Variants of this build are rolling out across LMIC tax authorities at different rates. Brazil and Chile are roughly a decade ahead. India, Indonesia, and Vietnam are in the middle. Most of sub-Saharan Africa, Pakistan, and Bangladesh are catching up fast. Singly, none of these data flows is a precision instrument for measuring a firm’s or an individual’s true liability. Collectively, they are the raw material from which such a measurement can be built.
The opportunity lies in integrating and inferring from the flows already being generated. The core task is linkage: the same firm or person surfaces across these datasets under different identifiers, and joining them needs a common key. Pakistan has one in the CNIC that NADRA maintains, which is what turns national-scale integration from aspiration into engineering.
The research literature has documented the underlying pattern cleanly. Carrillo, Pomeranz, and Singhal (2017) found in Ecuador that firms under-report line items that aren’t cross-checked by a third party. Pomeranz (2015) showed in Chile that VAT’s self-enforcement mechanism breaks down precisely where the third-party trail ends. Waseem (2017) estimates that 60–70% of Pakistan's potential VAT revenue is lost to under-reporting. None of these patterns is random. They are targeted, structural, and predictable: exactly the kind of thing supervised models can learn.
My own work with a subnational tax authority on a point-of-sale integration program in the hospitality sector provides a concrete figure for what reconciliation could deliver. Reconciling transaction-level device records with self-declared liabilities would have raised roughly 19% above the tax authority's actual collection from the sector. The broader finding of the paper is that digital visibility produces sustained compliance only when data are routinely translated into action (Asad and Cohen, 2026).
3. A framework: predict, target, enforce, learn
That missing reconciliation layer has a definite shape, and it is the same at every LMIC tax authority that has measured it: the enforcement pipe runs from a wide potential base at the top, narrows at every step, and reaches almost nothing at the bottom. For tax year 2024, Pakistan’s Active Taxpayer List crossed 7.27 million filers, a record. Most of those filers are individuals, salaried taxpayers, and small registrants, not the firms that generate the bulk of the liability. Against a 7.2-million-unit business universe, audits of any depth still run in the low tens of thousands. Modern risk targeting is an investment in widening the working end of that funnel.
The machine you build to attack a problem of that shape should be boringly procedural. Four steps, each of which is something machine learning is unusually good at, once the data is integrated.
Start with prediction.
The system begins at the densest part of the data, the registered base, and extends outward as registration grows. For each firm, a model estimates the liability it would report if it were honest, given everything observable about its activity, along with an uncertainty band around that estimate. That evidence includes signals of real activity (electricity and gas consumption, import volumes, point-of-sale and invoice feeds), the firm’s filing history, its ratios against sector and size peers, and its position in the supplier-buyer network that digital invoices reconstruct. The model is fit on audit-resolved cases, where realized under-reporting is observed, so it learns the empirical signature of evasion rather than a hand-coded rule.
Then targeting
Comparing each firm’s declared liability to its predicted range sorts the base: most fall inside the band and need nothing, while a thin tail falls well below it. Out of a few hundred thousand sales-tax-registered firms, perhaps tens of thousands sit in a high-risk zone and a few thousand at the very top. But a ranked list of suspects is only half the decision. For each case, the system also has an estimate of how much revenue is at stake and how clear-cut the evidence is, and those decide not just whether to act but how. The human attention is a scarce input, so the object is to spend as little of it on each case as the case truly requires.
Then enforcement
Most of the enforcement is not human work. Where the evidence is unambiguous and a declared figure simply contradicts an independent record, the discrepancy can drive an automated, evidence-backed action on its own: a notice showing the taxpayer the gap and the corrected figure, a demand, or a refund held pending reconciliation. Shown that the authority can already see the mismatch, many firms correct and pay before an officer ever opens the file, which also closes the encounter in which assessments are quietly bargained down. Cases escalate only as they get harder, from automated notice to desk review to field audit to investigation, and a scarce experienced auditor is spent only where the amounts are large, the structure is complex, or the evidence genuinely needs a judgment call. India’s DGGI unearthed substantial GST fraud in a single year, part of it in input tax credit fraud, caught by exactly this kind of automated reconciliation.
Finally, learning.
Each targeted case returns a label, whether it was settled by an automated notice or a full audit: revenue recovered or not, an assessment upheld or overturned on appeal. Feeding those labels back turns the system from a fixed rule into one that is continually re-estimated. The subtlety is that an audit has two kinds of value: the revenue it brings in now, and the information it adds to every prediction afterward. Always working the highest-risk cases maximizes the first and starves the second, since the model learns nothing about the firms it never looks at. So a share of capacity has to go to cases chosen for what they would reveal, the random sample among them, even when those cases are not the ones with the highest expected yield. That exploration is what stops the model’s picture of the economy from narrowing to the cases it already suspects, and it is why a system run this way keeps improving as its labeled history grows, while one that only exploits its current best guess plateaus.
None of these four steps is theoretical. Variants of all of them have run for years inside authorities that started earlier. The core methods are not new. What has shifted for the Pakistan-and-peers cohort is the cost of the conditions the framework needs: the administrative data has been digitized, and the price of storing, linking, and computing over it at a national scale has collapsed. The newest layer, foundation models, has less to do with the prediction itself than with the expensive work around it, resolving the same firm across inconsistent records and extracting structure from scanned returns and invoices. The capability is a decade old in rich countries; what is new is that it has become affordable for everyone else.
4. Invest, and evaluate
Two things follow for revenue authorities considering this kind of work.
First, the investment case can write itself. A national-scale prediction-and-targeting layer probably costs in the low single-digit millions of dollars per year to build and operate. The revenue at stake runs to hundreds of millions to billions per year in nearly every country in the LMIC cohort. The arithmetic is plain on its own, and only gets sharper as aid budgets shrink.
Second, and more importantly, these systems have to be continuously evaluated. They are scientific instruments. A model that ranked firms accurately in 2024 may rank them less accurately in 2026 as taxpayer behavior adapts. A targeting rule that worked in retail may fail in construction. Audit-selection algorithms, if not monitored, can drift toward geographic or demographic bias, eroding legitimacy faster than they raise revenue. The standard for evaluation here is the rest of empirical public finance: prespecified hypotheses, randomized pilots where the law permits, transparent reporting of results, and regular external review.
That evaluation discipline is what distinguishes serious tax-administration reform from a software contract. Several authorities, including Pakistan’s FBR, have already begun moving in this direction. That is the right instinct; the mistake would be to treat the tool as the finish line rather than the start of something to evaluate and learn from. The research literature has been converging on this template for years. Khan, Khwaja, and Olken’s work on Punjab property taxes ran as randomized trials. Best, Shah, and Waseem’s audit-deterrence study used FBR’s own randomized audit program. My own work with tax authorities follows the same template. The frontier in this area has moved past whether the technology can be built. The open question is whether revenue authorities will treat each enforcement deployment as an experiment whose results, including null and negative ones, change what gets deployed next.
Authorities that commit to both will lead the next decade of credible tax administration in the developing world. Those who buy the tools without the evaluation discipline will produce the next generation of unused dashboards.
Sources and further reading
Pakistan: official and administrative sources
FBR, Inspectors-IR (BS-16) seniority list (525 officers finalised, 2015); FBR/FPSC recruitment notification advertising 492 Inspector-IR posts (2021). Current strength on the order of 600.
FBR, Revenue Division Yearbook 2024–25. Total FBR collection Rs 11.74 trillion (~10.3% of GDP federal); direct taxes 45.7% of total; income-tax composition: withholding ~58%, advance tax ~33%, payments with returns ~4%, Collection on Demand ~5% (Rs 267 billion, +110% YoY). FBR’s briefing to the Prime Minister put federal tax-to-GDP at 10.6% (Profit by Pakistan Today, July 2025).
Government of Pakistan, Gazette notification on Inland Revenue Service cadre (2014): 1,229 officers BS-17 and above (BS-22: 2; BS-21: 47; BS-20: 170; BS-19: 255; BS-18: 377; BS-17: 378). FBR’s August 2024 HRMS PER circular [C.No.1(1)ERM/2024] lists active BS-19/18/17 counts closely tracking this baseline. The cadre has not been materially restructured.
Pakistan Bureau of Statistics, Economic Census 2023. First-ever digital economic census; 7.2 million economic units (2.7M retail; 825K services; 188K wholesale; 643K production; 23K factories). Coverage and totals reported in Dawn and The News International.
Pakistan: press reporting
Business Recorder, “Performance scheme: FBR’s low-level employees protest denial of benefits” (April 2025). Cited figure: total FBR sanctioned establishment of 28,000 posts. From a news report citing FBR establishment figures rather than from a primary FBR statistical publication; treated here as a credible but secondary cite for total FBR sanctioned strength.
Dawn, “FBR misses collection target by Rs429bn in eight months” (1 March 2026). FY 2025–26 8-month collection Rs 8.12 trillion.
Profit by Pakistan Today, “80% of Tier-1 retail branches disconnected from FBR POS system” (21 February 2026). Compliance figures behind the headline POS coverage.
Profit by Pakistan Today, “FBR expands POS network, nearly 36,000 retail and restaurant outlets now linked” (26 April 2026). Tier-1 retailer count and branch coverage as of April 2026.
Tax capacity and aid: international assessments
IMF, “Building Tax Capacity for Growth and Development” (Departmental Paper, October 2025) and IMF G20 Background Note on “Enhancing Domestic Revenue Mobilization through Strengthening Revenue Administration” (2025). LIDC revenue potential ~9 pp of GDP above current levels; 15% threshold for a functioning state.
OECD, Revenue Statistics in Asia and the Pacific 2025 (Pakistan country note). Pakistan tax-to-GDP 10.5% (2023); regional comparators.
OECD DAC, “A historic decline in foreign aid: Preliminary 2025 ODA data” (April 2026). Total ODA $174.3 billion in 2025, −23.1% in real terms; further decline projected for 2026.
World Bank, “Pakistan: From Swimming in Sand to High and Sustainable Growth” / Strengthening Government Revenues country diagnostic. Pakistan’s tax capacity estimated at ~22% of GDP.
Risk targeting in practice
Government of India / DGGI press notes and taxguru / The Probe reporting. FY 2024–25 ITC-fraud detections of ~Rs 58,772 crore across 15,283 cases; total GST evasion detected ~Rs 1.95 lakh crore.
IMF Technical Note, “Essential Analytics for Compliance Risk Management” (2024). The state of the practice on risk-targeting in revenue administration.
Research literature
Asad and Cohen (2026), “Technology without Teeth: Evidence from Voluntary Point-of-Sale Integration.” Working paper. Hospitality-sector ePOS integration at a provincial revenue authority; reconciliation gap on the order of PKR 460M (~19%) among adopters.
Bergeron, Tourek, Weigel (2024), “The State Capacity Ceiling on Tax Rates.” Econometrica. Returns to enforcement effort in the DRC.
Best, Shah and Waseem (2021), “Detection without Deterrence: Long-Run Effects of Tax Audit on Firm Behaviour.” Direct study of FBR’s randomised VAT-audit programme; documents that FBR could not audit more than 5% of VAT filers a year, and that audit detection in Pakistan has no observable deterrence effect on subsequent compliance.
Carrillo, Pomeranz, Singhal (2017), “Dodging the Taxman.” AEJ: Applied Economics. Firms under-report the line items not cross-checked by a third party.
Khan, Khwaja, Olken (2016), “Tax Farming Redux.” Quarterly Journal of Economics. Performance pay for property-tax collectors in Punjab.
Khan, Khwaja, Olken (2019), “Making Moves Matter.” American Economic Review. Merit-based postings of property-tax inspectors.
Pomeranz (2015), “No Taxation without Information.” American Economic Review. VAT self-enforcement breaks down where the third-party trail ends.
Waseem (2017, 2020). Pakistan VAT evasion estimates (60–70% of potential revenue lost to under-reporting).






Great insights. I believe for this to be successful, it is imperative to encourage a research-first, experimental approach within tax authorities, both internally and in collaboration with external partners.