This feature was joint-funded by Wayne State University Libraries, the Leventhal Map & Education Center at the Boston Public Library, and the American Geographical Society Library, at UW Milwaukee. A big thank you to these orgs for funding improvements that make the site better for everyone!
Over the latter part of the summer I was able to add a feature to OldInsuranceMaps.net that greatly enhances its mosaic-making capabilities, and is the logical continuation of the multimask interface.
Users can now queue the creation of single-file mosaics for volumes after the masking is complete, making it much easier to produce the full complement of distribution formats and web services. Links to all these formats, and to trigger the mosaicking processes, are now in the Mosaic > Derivatives section of each map’s summary page.
The creation of these mosaics is managed through a new jobs system that allows us to queue mosaics for creation without having to manually trigger them from the backend (which is what I was doing before).
Check out the documentation to jump straight into learning how to use it.
Why is this useful?
The multimasking concept only gets you so far. You can make contiguous masks for all content in a volume, but then you are still left with 100 individual layers, like this example of Detroit, Mich., 1897, vol. 2,:
It’s actually much more useful to combine those layers into a single file, making it easier to display and distribute the mosaic in new formats—like a single tile service endpoint for the whole volume, or as downloadable file for use in desktop GIS software.
And while I had already scripted these functions, before this update I would just wait for someone to email saying they had finished masking the layers and it was time to make a mosaic. So with a trio of organizations interested in making a lot more mosaicked exports for use in their own web applications, it was a good time to try out something new.
No one likes a long-running process…
Creating these files is not that complicated to do; a mask is applied to each layer via a virtual raster (VRT), and each of those is added to one final mosaic VRT and then exported to one big GeoTIFF. But, it can take many hours, and they can get pretty big; the one shown above is about 400mb but they could easily be 2-5 times that size, depending on the geographic extent of all the layers to be combined. So I made a fairly basic queuing system, using Celery, a package I was already using to handle background tasks like warping and cleaning up abandoned georeferencing sessions.
It works like this:
- The database now has a “jobs” table that stores configurations for each process.
- Each job has a status:
queued,running,completed, orerrored. - Users create new job entries as desired, and new jobs have the status of
queued. - The existing
maincelery worker now checks, every minute, forqueuedjobs, and starts them in the order they were queued (first in, first out). - A new
mosaiccelery worker is now dedicated to only running the actual mosaic processes, so when the main worker kicks the job off, its the mosaic worker that actually performs the task.- This leaves the main worker available to keep doing all of the georeferencing and other tasks needed throughout the day.
I was dubious about the feasibility of kicking off tasks in one worker from another, but it works great, actually. Again, more detailed info is available in the docs.
Also, because it is easy to swamp the server with multiple mosaic processes running at once, the current configuration will work through the queue uninterrupted but only allow two jobs to run concurrently.
Same system different formats
The appeal of setting this all up was not only to make it easier to generate mosaic COGs, but for the particular application we had in mind we also needed static XYZ tilesets that could be uploaded to an external storage bucket. So, once the job running framework was in place, I just slotted in a different function to build tilesets instead of COGs. In the future, we hope to add PMTiles output to the mix as well.
Room for improvement
There is actually a lot of variability from job to job, in terms of how long it will take and the size of the output file.
For example, as I write this, one particular job has been running for 16 hours! This is not ideal, and it’s due only to the fact that this particular volume covers peripheral parts of a Oklahoma City, so the layers are spread out across a really wide geographic extent, and there is a huge amount of blank space in between.
This means that the process takes forever, and the resulting GeoTIFF will be really, really large, and mostly empty on top of that :/ So, there are surely better approaches for edge cases like this, and maybe down the road we’ll add better ways to handle them. (A static XYZ tileset, or PMTiles file, would be a more efficient way to store this layer.)
It would be good to actually display the size of these files as well, so people would have a better idea of what they were about to generate, or, later, download.
Who can run these processes?
For now, there is a group of users with the ability to trigger mosaic jobs on any maps, but all users can do so on maps that they have “loaded”. So if you had requested maps of a city to work on, loaded it, have georeferenced everything (that you can figure out), and completed the multimask, you can go ahead and queue the creation of a mosaic for that volume.
Let us know how it goes! Join the OSM US Slack (introduce yourself in the #welcome channel!) and find us in the #oldinsurancemaps channel.