‘Putting data on a server does not make it statistically useful’

Independent surveys are important to check whether administrative data reflect reality, says P.C. Mohanan, former acting chairperson of the National Statistical Commission.

WrittenBy:Prachi Salve& ISignal
Date:
Article image

India is preparing to revamp parts of its official statistical system, conduct at least 14 surveys under the National Sample Survey Office this year, and make better use of its administrative and digital data. But, the Ministry of Statistics and Programme Implementation’s (MoSPI) budget for 2026-27 is 22 percent below its estimated expenses, and will not be enough to meet the physical targets of its programmes, the Ministry told a Parliamentary standing committee in a report published in March 2026. On the other hand, it had surrendered Rs 6,088 crore ($711 million) over the last three financial years.

For 2026-27, MoSPI estimated expenses of Rs 5,826 crore ($625 million), but allocations in the Union budget amounted to Rs 4,522 crore ($485 million) for schemes, a shortfall of about 22 percent. Of the amount allocated, 87 percent is for the Member of Parliament Local Area Development Scheme (MPLAD), under which each MP is provided Rs 5 crore each year to support local developmental activities.

The Ministry told the Committee that against an estimate of Rs 777 crore—or Rs 1 crore per district—under the Statistical Strengthening Sub-Scheme, revamped to help states generate good quality data at sub-State level, allocations stood at Rs 15 crore, based on past expenditure. The Ministry’s estimate itself was about half the requirements of Rs 1,400 crore as indicated by states and Union territories, it said.

In addition, the National Sample Survey aims to conduct 14-15 surveys this year at an estimated budget of Rs 788 crore, but allocations amounted to Rs 415 crore.

But money is only one part of the problem.

In December 2025, the Committee's review of the National Statistical Commission had raised questions about the institution's ability to perform its oversight role. The NSC has a mandate to monitor and enforce statistical standards, but it remains a part-time body with a secretariat supported by MoSPI.

How serious are these institutional and technical weaknesses, and can India's statistical system keep pace with the country’s ambitions for more granular, digital and data-driven policymaking? To understand this, we spoke to P.C. Mohanan, former acting chairperson of the National Statistical Commission (NSC).

Mohanan has been a government statistician for over 30 years and was a member of the NSC from June 2017 till his resignation in 2019. He was involved in unemployment, migration and healthcare surveys at the National Sample Survey Office (NSSO). Mohanan was a member of the expert committee on agricultural statistics, estimating housing shortage, and an evaluation committee appointed after the 2006 Sachar report, which studied the condition of Muslims.

Edited excerpts:

The NSC is ​​mandated to perform statistical audits, but “its only notable action was an audit of the Index of Industrial Production in 2011”, the Parliamentary Standing Committee on Finance noted in December 2025. How effectively has the NSC been able to perform its oversight role?

The oversight role of the NSC is very limited. The Commission is a part-time body that meets perhaps once in two months, and most of its members and the Chairperson have other engagements. The time and resources they can devote exclusively to the Commission are therefore restricted.

The Commission has to depend on a very small secretariat. When I was part of the first NSC secretariat, from 2006 to 2009, I was heading the Secretariat, along with another officer supporting. However, the first Commission was able to undertake several important initiatives. For example, it took a firm position on bringing together the various price indices into a uniform rural, urban and combined index, and there were also several changes in the National Sample Survey system.

I have not seen the Commission taking similar major initiatives to change or reorganise the statistical system in subsequent years.

The problem is also that, unlike commissions with full-time members and their own expert staff, the NSC has limited resources at its command. In many cases, the Ministry or MoSPI decides the agenda or brings specific questions before the Commission. The Commission discusses them and makes recommendations, but how deeply it can examine the issues is doubtful because it has to depend on presentations from the same officials who are involved in producing the statistics.

So, to a great extent, the oversight role is absent at present. The Commission functions more through the academic standing and influence of its Chairperson and members than through institutional authority.

The government is pursuing several reforms to improve India's statistical system, including changes to surveys, data standards and the use of administrative and digital data. In this context, the Committee flagged the lack of statutory status to the NSC, as envisaged by the Rangarajan Commission. Do you think these reforms can improve the credibility of official statistics without a more independent NSC, or is institutional reform also necessary?

It is encouraging that several reforms have been initiated in the last two years to update key economic data. We also now have a good data ecosystem, with many data journalists examining and questioning government data. To some extent, this provides an independent check.

But this is not enough. Most statistical processes are still not transparent or sufficiently understood by the media and researchers to allow them to critically examine the numbers.

An independent oversight body has an important role, particularly when new statistical initiatives are introduced. A good example is the last Economic Census. When I was a member of the NSC around 2017-18, we had reservations about conducting another Economic Census because of concerns about the quality of the data and the fact that the results of earlier censuses had not been used effectively.

We had said that another Economic Census should not be undertaken unless these issues were addressed. However, the Ministry went ahead, obtained Cabinet approval and entered into an agreement with the Common Services Centres for the fieldwork. The NSC was not involved in that decision-making process.

The fieldwork was ultimately carried out through locally recruited personnel, and in my assessment the quality of the resulting data was extremely poor. I understand that the expenditure was around Rs 1,000 crore [Rs 913 crore allocated for the seventh economic census], but I believe the exercise did not produce results of sufficient value.

The lesson is that such a large statistical exercise needs to be technically guided and its design and implementation should be subject to serious scrutiny. Before approving major surveys and censuses, the views of the NSC should perhaps be mandatory. Some statutory support for the Commission is therefore necessary.

At the same time, in the present context of increasing centralisation, giving more powers to the NSC through statutory backing could also be uncomfortable for the government.

MPLADS accounts for nearly Rs 3,952 crore of MoSPI's Rs 5,503 crore budget in 2026-27, even as the Ministry says that its allocation for statistical programmes is not enough to meet its targets. Since MPLADS is very different from the Ministry's core statistical work, does having the scheme within MoSPI make it harder to understand how much the government is actually spending on statistics?

The MPLADS budget is more or less fixed because each MP receives a fixed amount for constituency projects. So, to that extent, it cannot really be mixed with the funds meant for statistical activities.

However, I am not clear why the government has clubbed MPLADS with statistics, because the two have completely different functional specialisations. This gives a wrong impression about the budgetary support available for statistics.

More than scaling up funding, I think there is an urgent need to strengthen human resources.

Most of the new data-collection activities of the NSSO are being carried out through contract staff. Field investigators require considerable training to understand sampling procedures and statistical definitions, but contractual investigators may be engaged only for six months or a year. This affects the retention of skills and potentially the quality of the data collected.

The field staff strength has not kept pace with the growth of the country's population or the expansion of surveys. There is a need to increase lower-level field staff and have more permanent personnel so that they gain experience over time.

The same problem exists with the Indian Statistical Service. Its overall strength has remained more or less unchanged for the last 30 to 40 years. There has been an increase in higher-level posts, but this has often come at the cost of lower-level positions. You need people not only at the senior level but also at the lower levels to collect, process and analyse the data.

Under the Statistical Strengthening Sub-Scheme, the government is changing the way it funds statistical work in the states by linking allocations to outputs and milestones. Given that states play an important role in producing India's official statistics, do you think this change will improve the quality of data coming from the states?

The national statistical system also includes state statistics, but most states have not been able to reorganise their statistical activities in response to changing technology and data requirements. There is considerable variation between states, and there is also not enough clarity about the statistical outputs expected from them.

The earlier scheme for strengthening state statistical systems did not result in any sustained improvement. Having recently headed a state-level statistical commission, I can say that it is not easy to bring about uniform changes because the problems are different in different states. MoSPI needs a much stronger team to address these state-level issues.

One important problem is the position of the Directorate of Economics and Statistics within the state government. The director is often relatively low in the administrative hierarchy and does not have strong administrative powers. Statistical work is also decentralised across different departments such as industry, employment and finance. So the statistics directorate has limited influence over decisions and often has to obtain approvals from finance departments or other departmental secretaries.

There is also considerable variation in capacity. Some states have strong survey units, while others do not. Even in the case of State GDP, which is one of the most important indicators for state governments, many states do not have enough specialised personnel. People move in and out of these positions, and the necessary expertise is not always retained.

The government has tried to address this through schemes such as the India Statistical Strengthening Project [now the Statistical Strengthening Sub-Scheme], under which states received funding for statistical activities. But this did not bring about sustained improvement.

Another problem is that there is much greater interest in national-level statistics than in state-level statistics. The media and researchers closely follow national data, but there is much less scrutiny and use of state-level statistics. So the lack of demand for high-quality state statistics is also part of the problem.

The Standing Committee has pointed out that government data is spread across different ministries and agencies and is often collected and stored in different ways. Does this make it harder to produce timely and reliable official statistics, and what needs to change to make better use of the data the government already has?

The Standing Committee is absolutely right to highlight this. Different metadata standards, differences in data coverage and operational issues mean that a lot of government data exists in silos that cannot easily communicate with one another. Data integration has always been a challenge in official statistics.

The problem is particularly acute with administrative data. [These are data collected and maintained by official agencies, such as when enterprises file various periodical returns, traders provide data to the customs, tax payees file returns, etc.]

Different government agencies collect data using their own methodologies, whether through statutory returns [data collected through mandatory filings] or beneficiary enrolment. Digitisation means that this information can now be collected through apps and stored on central servers, but simply putting data on a server or displaying it on a live dashboard does not make it statistically useful.

The coverage and reliability of administrative data are often unclear. For example, when industrial classifications are updated, existing administrative datasets may not be updated accordingly. Similarly, if an industrial unit changes its product mix, this may not be reflected in the basic information held by the department. This can affect corporate filings, including MCA21 data [e-governance portal to register businesses, file statutory annual returns, manage director details, and make regulatory payments].

One major problem with administrative data is that we often don't know whether its coverage has remained consistent over time. Take small industries: their definitions and classifications can change. If the definition changes, you cannot necessarily compare this year's data with data from five years ago.

This is different from a properly designed sample survey, where you know the population being covered and can assess whether the sample is representative.

There are many examples of administrative databases where this is a problem. Employment exchange data, for instance, was once expected to form the backbone of unemployment statistics, but people who found jobs often did not remove themselves from the registers, while many unemployed people never registered in the first place.

The same issue can arise with health data. If you use the Health Management Information System, for example, you need to know which hospitals are actually covered. If many hospitals are outside the system, you cannot simply interpret the numbers as representing the entire population.

This is also why I have reservations about treating live dashboards as statistical evidence. Take the e-Shram portal: the government reports more than 31 crore registrations, but that number does not tell us the total number of unorganised workers. We don't know the coverage, how many eligible people have not registered, or how many people have subsequently ceased to be eligible but remain in the database.

Similar issues arise with dashboards for schemes such as Swachh Bharat or Jal Jeevan Mission. The dashboard may report very high levels of achievement, but those figures need independent verification.

Independent surveys are therefore important for checking whether administrative data reflects reality.

India is changing the way it calculates GDP, including moving the base year from 2011-12 to 2022-23 and bringing in more government data and newer surveys. What do you expect this change to improve, and what should people be looking out for when the new GDP numbers are released?

Changing the base year is a basic requirement to incorporate new datasets, capture emerging sectors and meet international standards. These changes make GDP a more relevant reflection of the economy.

The new series contains several methodological improvements. For example, the share of services in the economy has come down and the share of industry has gone up because the use of MCA data has changed the way some activities are classified. The new series also introduces double deflation, where inputs and outputs are deflated separately to calculate constant-price estimates. These are technically sound improvements.

But the new series is not strictly comparable with the previous base-year series. One has to be particularly careful when looking at the back series because some of the new datasets used for the new GDP estimates are not available for earlier years.

There is another important issue at the state level. Ideally, State GDP estimates should be calculated independently and then added together to arrive at national GDP. In practice, national GDP is calculated first and State GDP estimates have to be consistent with that national total.

For example, some economic activities such as railways are calculated at the national level and then apportioned among states using parameters such as the length of railway lines. The problem is that this means the estimates are not always built up from the state level to the national level; in some cases, they are allocated down from the national level.

This becomes even more difficult when states want district-level GDP estimates, because much of the data used at the national level is not available at the district level. So there are still significant limitations in producing reliable district-level estimates.

Overall, I would say the new national GDP methodology is an improvement over the previous one. The earlier series had several areas where it was not clear how particular estimates had been arrived at, whereas the new methodology is somewhat more transparent. But users should be cautious when making long-term comparisons across the old and new series.

And as administrative data becomes increasingly important for GDP, it should not simply be accepted without independent verification. Administrative data has the advantage of being generated regularly and, because of digitisation, becoming available much faster. But its reliability and comparability over time need to be independently checked through surveys or other methods. Otherwise, changes in the administrative system itself can be mistaken for changes in the economy.

This interview is republished with permission from ISignal, formerly IndiaSpend, a data-driven, public-interest journalism non-profit. It has been lightly edited for style and clarity.

Also see
article imageIn Bengaluru’s water-stressed areas sit most of its data centres

Comments

We take comments from subscribers only!  Subscribe now to post comments! 
Already a subscriber?  Login


You may also like