{
  "is": "issue",
  "title": "Emergency MariaDB maintenance for opal4, opal7, and opal9",
  "body": "\u003cp\u003e\u003cem\u003eFixed\u003c/em\u003e - The MariaDB recovery on opal9 is complete.  Databases with names starting with \u003ccode\u003etro\u003c/code\u003e through \u003ccode\u003ezzz\u003c/code\u003e are being restored from our 2 December 2023 backup.\u003c/p\u003e\n\u003cp\u003eOur plan now is to test the recovery process on a test server to determine why we were not able to import the dump files and then refine our maintenance procedures based on those findings. \n  \u003cspan class=\"faded\"\u003e(22:21 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - The MariaDB recovery on opal9 is still in progress. We expect the recovery to be complete in around 90 minutes. \n  \u003cspan class=\"faded\"\u003e(21:21 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - The MariaDB recovery on opal7 is complete.\u003c/p\u003e\n\u003cp\u003eThe MariaDB recovery on opal9 is still in progress. We\u0026rsquo;ll update with an ETA as soon as possible. \n  \u003cspan class=\"faded\"\u003e(21:12 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - The MariaDB recovery on opal7 and opal9 is still in progress. We expect that it will take at least a couple of hours more to complete. \n  \u003cspan class=\"faded\"\u003e(19:39 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - The MariaDB recovery on opal7 and opal9 is still in progress. We expect that it will take at least a few more hours to complete.\u003c/p\u003e\n\u003cp\u003eThe MariaDB recovery on opal4 is now complete, however we were not able to use the most recent backup for about 50% of the recovered databases. The remaining databases were recovered as follows:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eDatabases with names starting with \u003ccode\u003elsa\u003c/code\u003e through \u003ccode\u003eshn\u003c/code\u003e were restored from our 2 December 2023 backup.\u003c/li\u003e\n\u003cli\u003eDatabases with names starting with \u003ccode\u003esho\u003c/code\u003e through \u003ccode\u003ezzz\u003c/code\u003e were restored from our 1 December 2023 backup. \n  \u003cspan class=\"faded\"\u003e(16:56 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - It seems some opal4 databases were not restored. We\u0026rsquo;re restoring those databases now.\u003c/p\u003e\n\u003cp\u003eThe recovery on opal7 and opal9 is still in progress. \n  \u003cspan class=\"faded\"\u003e(14:55 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - The MariaDB recovery on opal4 is complete. Recovery on opal7 and opal9 is still in progress. \n  \u003cspan class=\"faded\"\u003e(14:33 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - The database restore is still in progress. We expect that it will take at least a few more hours. \n  \u003cspan class=\"faded\"\u003e(13:47 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - What happened?\u003c/p\u003e\n\u003cp\u003eWe took the MariaDB database down to stop the ibdatafile from growing exponentially. The only way to do this is to dump all databases, delete the log files, update the server configuration, and restore the data.\u003c/p\u003e\n\u003cp\u003eThe dump process worked without any errors. The clean up and server configuration completed without any errors. The restore process is where things failed.\u003c/p\u003e\n\u003cp\u003eBecause of errors in 1 or more database(s) transactions failed to complete causing the restore to never get past a certain point. The amount of data in the mysql data directory would increase up to a certain point and then the data size would drop by half before repeating. After running the restore process twice and having errors in different spots we decided to extract each database from the monolithic backup we had taken previously. Why a monolithic backup? It\u0026rsquo;s usually faster to dump and restore, except in this case.\u003c/p\u003e\n\u003cp\u003eThe extraction process is painfully slow compared to just running a working dump restore. That\u0026rsquo;s why this process is taking so long.\u003c/p\u003e\n\u003cp\u003eWhat didn\u0026rsquo;t we do?\u003c/p\u003e\n\u003cp\u003eWe could have restored from the latest backup which would have been 24 hours or less old, however, that option comes with the significant risk of data loss. Rather than risk losing 24 hours of data we went with the slower, safer process.\u003c/p\u003e\n\u003cp\u003eThe current restore process for each server is now past the previous failure point.\u003c/p\u003e\n\u003cp\u003eWe\u0026rsquo;d sincerely apologize for the downtime this has caused for your apps that use MariaDB. In the tests we performed before the actual event we did not run into any of these errors. \n  \u003cspan class=\"faded\"\u003e(12:15 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - The data restoration process is still ongoing. There is currently no ETA we can provide. \n  \u003cspan class=\"faded\"\u003e(10:29 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eWatching\u003c/em\u003e - Opal4, Opal7, Opal9: We are still restoring data to the MariaDB databases. The exact amount of time left in the restoration process is unknown but we will update every hour until all of the restores have finished. \n  \u003cspan class=\"faded\"\u003e(08:42 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n\u003cp\u003eOn Sunday, 03 December 2023 at 0500 UTC we\u0026rsquo;ll be taking the managed MariaDB database service offline for maintenance on the following servers:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eopal4.opalstack.com (Dallas)\u003c/li\u003e\n\u003cli\u003eopal7.opalstack.com (Phoenix)\u003c/li\u003e\n\u003cli\u003eopal9.opalstack.com (Frankfurt)\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThe maintenance window is 2 hours. During the maintenance sites and applications which use MariaDB (including WordPress sites) will not function.\u003c/p\u003e\n\u003cp\u003eWe apologize for the short notice and any inconvenience. If you have any questions or concerns regarding the maintenance then please contact our support team. \n  \u003cspan class=\"faded\"\u003e(07:23 UTC — Dec 3)\u003c/span\u003e\n\n\u003c/p\u003e\n",
  "createdAt": "2023-12-03 07:23:00 +0000 UTC",
  "lastMod": "2023-12-03 07:23:00 +0000 UTC",
  "permalink": "https://opalstackstatus.com/issues/2023-12-03-emergency-mariadb-maintenance-for-opal4-opal7-and-opal9/",
  "severity": "down",
  "resolved": true,
  "informational": false,
  "resolvedAt": "2023-12-03 22:21:55",
  "affected": ["Shared: Dallas (USA)", "Shared: Frankfurt (DE)", "Shared: Phoenix (USA)"],
  "filename": "2023-12-03-emergency-mariadb-maintenance-for-opal4-opal7-and-opal9.md"
}