My first venture into AI has been fun all the way.
I now have a 6,720 line Fortran source code which is one component of what makes my current project achievable.
It runs on my oldest computer, with a Windows 98 system, the last system to support my ancient Fortran compiler.
It's design function is to take the shaggy output of my OCR software and manipulate it into a format that Windows
EXCEL finds edible (lines of comma delimited text data).
The "Hi Tech" production sequence of my current project is:
1) Record, in handwritten format, details of a marriage in the city of Palermo during the 19th century.
(Done by others, mainly priests, over the course of most of a century,
2) Record, in standardised typed format, summary extracts from the original handwritten documents.
(Done by others, government employees, over the course of about 20 years, around 1885-1905).
3) Create high quality images of the old type set pages.
(Done by others, likely within the last 20 years)
4) Digitize the image set. (One complete image set, plus parts of a second image set of a
second typed document printing.) Done by me, taking a few months.
5) Process the digitized images into ragged text files using OmnipagePro 14 OCR software.
First production run complete. Second production run, with added error correction, still in process.
6) Transform the ragged text files into well formatted text files. Using AI software.
Writing software, about 2 months. Apply the AI software, about 10 seconds per 100 text pages.
7) Transform the formatted text files into elegant EXCEL tables, using standard Microsoft software.
Quick, any size upto about 70,000 lines of information. These tables are sortable, excellent feature.

About the only format change between the original typed pages and the EXCEL table is that the date
has been subdivided into three columns to allow easy time sorting.
Future plan is to sort the whole data set from a marriage table into a family groups table. This sorting
will be easily achieved in EXCEL, but requiires higher data accuracy in the columns used for the sorting.

The data set is a good 200,000 lines long. Contains over 100,000 marriages (double listed) and around
600,000 persons full names, six per marriage in most cases.

It is published on two sites, the first leads to the second. I love the EXCEL files in the second site. They
are too large for the 1 meg file size restriction of the first site. Both sites seem stable and are free.

http://freepages.genealogy.rootsweb.ancestry.com/~tornabene/ see "Palermo Marriage Index"