Data Insight: Automatic Coding of Occupations: Methods to create the Scottish Historic Population Database (SHPD)

Read publication

What we found

Preliminary experiments have been undertaken using a relatively small pilot dataset (birth records from 1900 and marriage records from 1930) and obtained reasonable results from a combination of exact matching and statistical classification.

In the pilot, by combining exact matching for texts that have been seen in the training data and the Bayes classifier for the rest, the accuracy levels achieved from cross-validation are 94-97%.

Currently, experiments using a larger section of coded training data (50,000 occupations) have uncovered a number of interesting features:

  • Because some occupations are very common, a relatively small set of manually coded strings covers a very large proportion of the SHPD records. In applying these 50,000 differ ent randomly-selected strings we have coded 22 million of the 25 million records (almost 88%). At time of writing, not all 31 million were available to process.
  • A significant number of the records that are not exact matches contain spelling errors or illegible letters.
  • Some records contain more than one occupation, particularly in war years when it is common for both a civilian and a military occupation to be recorded. Dividing these into the individual occupations would allow more to be matched exactly.
  • The military occupations often contain details such as a regiment which is irrelevant to the rather coarse HISCO coding, and this usually prevents them from being exactly matched.
  • Present work includes CDDA undertaking a review of coded occupations, part of which was to review spelling in the occupation strings coded. This in turn provides a large corpus of correctly-spelled, relevant words that can be used to detect and correct errors in the un-coded data. Examples of possible corrections are: BULLDER to BUILDER, FISHRMAN to FISHERMAN, FARMA to FARM, and PLOUGHMNA to PLOUGHMAN).

Why it matters

By completing both current and future work, it is envisaged that this will allow well over 90% of the occupations to be coded as exact matches. Those strings that have not been matched will be passed to a statistical classifier trained on the manually coded records. Excitingly, once occupations are fully coded, this will be a fantastic resource for researchers to access to some 23 million individuals dating back to 1856. Finally, for the first time the UK will have a data system of similar depth and breadth to those in Scandinavia.

Share this: