@VirMach said: I agree with you on this. We didn't really have another choice for these. It's painful and bad, for us as well. There are probably 5% of people that have been stuck for 48 hours now and I don't like that but we're doing all we physically can.
Just in case you're thinking this is about the migrations, the discussion you're referring back to is actually about Ryzen Location Change button and @yoursunny opinions on it.
@VirMach said:
Unfortunately there's pretty much nothing we can do about that really. I really think SolusVM did some updates to that tool and broke it for older operating systems recently. They've been breaking a lot of things, like libvirtd incompatibility, the migration tool wasn't working for a while, the operating systems don't template properly and they haven't been syncing properly a lot of times. They've just been updating all their PHP versions and whatever else, racing forward without actually checking anything.
Ok, I understand that, but is there a way to fix network manually? What settings should i set to get vm vorking? I couldn't see any problems from the first approach - there is eth0 ( ens3) interface up, there is a route to gateway xxx.xxx.xxx.1
I can ping nearby nodes, but for gateway response is Network host unreachable.
@Mumbly said: @VirMach that's misunderstanding now
I didn't comment or criticized Virmach migrations here, but discussed @yoursunny's suggestion (well, he made a few good suggestions but this one I feel like wasn't the best one) about 24 hours migration queve in the future
I understand, I'm just adding onto it that I agree with you and it's not user friendly and I'm also saying that unfortunately our current situation is similar and also not how we intended for it to be coded by the developer. I did read what @yoursunny suggested and that definitely has its benefits but our ideal version would be immediate.
I remember we improved our script to minimize downtime to only something like an hour per VM and had that concept working pretty well, but it didn't work out when you're planning for efficiency and many servers.
How the project was for the developer that never completed it:
Customer is eligible to use it when the server is queued for migration, and later on it would be open to everyone (the latter being the Ryzen to Ryzen idea.)
Customer receives one credit per term length, as in if they could cancel and re-order, then they're eligible for it. Therefore, naturally customer on first month of service wouldn't be able to abuse it.
Migration is queued in a batch, but the batch gets processed more immediately. It essentially would only wait for the right load conditions and be throttled by quantities of requests. If there are instantly 1,000 requests, then it could theoretically naturally take 24 hours.
Since AFAIK SolusVM does not have an API to run the script that completes a migration, AKA, marks it in the database as being on a new node, our idea was that it'd power up a new service, then replace the details of the existing service on WHMCS.
Old service gets marked for deletion but not immediate deletion in case anything goes wrong, so the data is still there for a few days. More aggressive pruning may happen during peak usage, it'd essentially have let's say a 200GB pool where it stays for a week, and anything past that maybe only a day. This also allows for an easy revert button to be worked in later in customer regrets his decision, instead of contacting us and instead of moving it twice.
It'd have been pretty nice if we had that ready in time for these migrations so it instead didn't end up as the vague day-long periods.
Trying to figure out from network status and backlog here: Should FFME002 be operational?
Have never gotten it to work here ...
Troubleshooter reports:
Main IP pings: false
Node Online: false
Service online: offline
Operating System: linux-ubuntu-16.04-server-x86_64-minimal-latest
Service Status:Active
Something strange about FFME003 network config of my vm. If i set network interface to dhcp, i receive correct ip address (xxx.xxx.163.xxx) but from dhcp server of another subnet xxx.xxx.162.2. And ip address i receive is from the same subnet as FFME004 vm. Is this working as intended? If i restore network config from control panel, i get the same ip as static, excluding missing symlink to resolvconf.
Customer receives one credit per term length, as in if they could cancel and re-order, then they're eligible for it. Therefore, naturally customer on first month of service wouldn't be able to abuse it.
Most services are paid annually.
One migration per year is not enough.
That's why I suggested once per month.
Some starter credits should be granted on the current services, because:
Some services have been auto-migrated to undesirable locations, such as Amsterdam to Frankfurt.
Looking glass nodes are inoperable, so that for all the migration done so far, the chosen location may be unsuitable for the needs.
Once all the looking glass nodes are up, can we at least have 2~3 credits per year?
That's a lot more flexible than only one credit per year, in case the network condition deteriorates in the middle.
We accept Karma donations for the last flan. 🍮 affbrr
@realEthanZou said:
Seems many JP VMs got their IP changed without prior notice
Indeed. No email, and no information I could find on the Network Status page.
While I understand that a host can face many challenges that cannot be predicted or immediately explained, this is not one of such situations. And the lack of a warning and the lack of information about why the change happened (Was it intentional and is it to stay, or was was it just a configuration mistake that will be reverted?) is a big fat NO NO in my book.
What tha F is going in with FFME? Migrated from FFME004 yesterday for 3 bucks fee, got the same FFME004, but online and working, today it's totally offline - no boot, no VNC, nothing. And no information about what's going on. I was waiting patiently for two weeks, created zero tickets, but now i have to give up and leave as soon as i could get my data from FFME003. Honestly, even my home server with my hobbyist approach does have less downtime and more reliability, and migrates faster.
While our Control Panel still shows 0 IPv4 addresses, we can finally get back in at the assigned IPv4 address and also through our private VPN. So progress has been made at least on the RYZE.SEA-Z002.VMS node in Seattle. We'll see if it stays up.
We're ramping up the abuse script temporarily. If you're at around full CPU usage on your VPS for more than 2 hours, it'll be powered down. We have too many VMs getting stuck on OS boot and negatively affecting others at this time, due to the operating systems getting stuck on boot for some after the Ryzen change. I apologize in advance for any false positives but I do want to note that it's technically within our terms for 2 hours of that type of usage to be considered abuse, we just usually try to be much more lenient.
@realEthanZou said:
Seems many JP VMs got their IP changed without prior notice
@realEthanZou said:
Seems many JP VMs got their IP changed without prior notice
Indeed. No email, and no information I could find on the Network Status page.
While I understand that a host can face many challenges that cannot be predicted or immediately explained, this is not one of such situations. And the lack of a warning and the lack of information about why the change happened (Was it intentional and is it to stay, or was was it just a configuration mistake that will be reverted?) is a big fat NO NO in my book.
We did send out emails, but it's possible they did not all send due to SolusVM also being overloaded around that time. Please also check your spam box, we sent these directly from SolusVM.
Customer receives one credit per term length, as in if they could cancel and re-order, then they're eligible for it. Therefore, naturally customer on first month of service wouldn't be able to abuse it.
Most services are paid annually.
One migration per year is not enough.
That's why I suggested once per month.
Some starter credits should be granted on the current services, because:
Some services have been auto-migrated to undesirable locations, such as Amsterdam to Frankfurt.
Looking glass nodes are inoperable, so that for all the migration done so far, the chosen location may be unsuitable for the needs.
Once all the looking glass nodes are up, can we at least have 2~3 credits per year?
That's a lot more flexible than only one credit per year, in case the network condition deteriorates in the middle.
Well initially the way it's going to work is there will be a period of time where you're allowed to "Ryzen Migrate" to your desired location (without data) as everyone lands in their desired location. It'll be announced here and on OGF, as well as most likely an "Announcement" on our website and probably a 1-2 week period where this can be done by everyone eligible.
The credits I described are for a system not yet coded.
@Papa said:
What tha F is going in with FFME? Migrated from FFME004 yesterday for 3 bucks fee, got the same FFME004, but online and working, today it's totally offline - no boot, no VNC, nothing. And no information about what's going on. I was waiting patiently for two weeks, created zero tickets, but now i have to give up and leave as soon as i could get my data from FFME003. Honestly, even my home server with my hobbyist approach does have less downtime and more reliability, and migrates faster.
FFME004, we found ECC error, and the settings also dropped off again. Memory swap fixed FFME004, we couldn't send out migration emails in time because xTom worked very quickly to get this replaced. The setting drop-off caused a disk to drop and node was online but VMs were not booting. That has been resolved. I'm checking it again to see if settings stick, if they don't there might be another reboot which we'll create network status for but hopefully this will be stable moving forward.
FFME has had 3 network status, many updates here, on OGF, and probably a fair share of emails. It can be considered an ongoing issue until they prove themselves by remaining online for more than 2 days.
Settings stuck on FFME004 but I'm pretty sure I've said that once before. There's zero information on this but there's been constant kernel bugs regarding these fixes on Linux. I can't rewrite the Linux kernel right now so until Linux figures out how it's going to treat these problems I don't know what else to do about it.
These issues have come on gone ever since NVMe SSDs existed, you can search online since around 2015. If you just look at Linux bug trackers it seems like every version fixes one thing and breaks another. I'll have to come up with some kind of kernel update plan moving forward where we try to mitigate the issues from re-appearing. But of course the solution can't just be to stay on the same version for 5 years.
We're using literal copies of the same nodes in FFM... in Tokyo, and they're not having the same problems with the only difference being kernel versions.
@NerdUno said:
While our Control Panel still shows 0 IPv4 addresses, we can finally get back in at the assigned IPv4 address and also through our private VPN. So progress has been made at least on the RYZE.SEA-Z002.VMS node in Seattle. We'll see if it stays up.
Spoke too soon. Status back to dead in the water this afternoon.
@VirMach said:
Settings stuck on FFME004 but I'm pretty sure I've said that once before. There's zero information on this but there's been constant kernel bugs regarding these fixes on Linux. I can't rewrite the Linux kernel right now so until Linux figures out how it's going to treat these problems I don't know what else to do about it.
These issues have come on gone ever since NVMe SSDs existed, you can search online since around 2015. If you just look at Linux bug trackers it seems like every version fixes one thing and breaks another. I'll have to come up with some kind of kernel update plan moving forward where we try to mitigate the issues from re-appearing. But of course the solution can't just be to stay on the same version for 5 years.
We're using literal copies of the same nodes in FFM... in Tokyo, and they're not having the same problems with the only difference being kernel versions.
My VM in FFME004 seems fine now, but I have a couple of VMs FFME005 and FFME006 having Status "Offline", can't bootup, can't reinstall OS. Hope you can help, thanks! Ticket #754039
@VirMach said:
We're ramping up the abuse script temporarily. If you're at around full CPU usage on your VPS for more than 2 hours, it'll be powered down. We have too many VMs getting stuck on OS boot and negatively affecting others at this time, due to the operating systems getting stuck on boot for some after the Ryzen change. I apologize in advance for any false positives but I do want to note that it's technically within our terms for 2 hours of that type of usage to be considered abuse, we just usually try to be much more lenient.
Boot loop after migrating to a different CPU or changing to a different IP is not abuse.
Customer purchased service on a specific CPU and a specific IP that are not expected to change.
The kernel and userland could have been compiled with -march=native so that it would not start on any other CPU.
The services could have been configured to bind to a specific IP, which would cause service restart loop if the IP disappeared.
The safest way is not automatically powering on the service after the migration.
The customer needs to press Power On button themselves and then fixes the machine right away.
Running -march=native code on an unsupported CPU triggers undefined behavior.
Undefined behavior means anything could happen, such as pink unicorn appearing in VirMach offices, @deank stopping to believe in the end, or @FrankZ receiving 1000 free servers.
The simple act of automatic powering on a migrated server could cause these severe consequences and you don't want that.
We accept Karma donations for the last flan. 🍮 affbrr
@NerdUno said:
While our Control Panel still shows 0 IPv4 addresses, we can finally get back in at the assigned IPv4 address and also through our private VPN. So progress has been made at least on the RYZE.SEA-Z002.VMS node in Seattle. We'll see if it stays up.
Spoke too soon. Status back to dead in the water this afternoon.
This has an issue with a software getting stuck and duplicating its process over and over until it overloads and we have to reboot it. We made some changes, if it happens again we'll try to catch it earlier this time to avoid a reboot.
@VirMach said:
We're ramping up the abuse script temporarily. If you're at around full CPU usage on your VPS for more than 2 hours, it'll be powered down. We have too many VMs getting stuck on OS boot and negatively affecting others at this time, due to the operating systems getting stuck on boot for some after the Ryzen change. I apologize in advance for any false positives but I do want to note that it's technically within our terms for 2 hours of that type of usage to be considered abuse, we just usually try to be much more lenient.
Boot loop after migrating to a different CPU or changing to a different IP is not abuse.
Customer purchased service on a specific CPU and a specific IP that are not expected to change.
The kernel and userland could have been compiled with -march=native so that it would not start on any other CPU.
The services could have been configured to bind to a specific IP, which would cause service restart loop if the IP disappeared.
The safest way is not automatically powering on the service after the migration.
The customer needs to press Power On button themselves and then fixes the machine right away.
Running -march=native code on an unsupported CPU triggers undefined behavior.
Undefined behavior means anything could happen, such as pink unicorn appearing in VirMach offices, @deank stopping to believe in the end, or @FrankZ receiving 1000 free servers.
The simple act of automatic powering on a migrated server could cause these severe consequences and you don't want that.
We're ramping up the abuse script. It's what it is called. I didn't say boot loop after migrating is abuse.
Abuse script will just power it down, not suspend. I don't see the harm in powering down something stuck in a boot loop. I was just providing this as a PSA for anyone reading who might be doing something else not related that's also using a lot of CPU and for general transparency, we're making the abuse script more strict to try to power down the ones stuck in the boot loop automatically more quickly.
The safest way is not automatically powering on the service after the migration.
The customer needs to press Power On button themselves and then fixes the machine right away.
Not possible, we have to power up all of them to fix other issues. Otherwise we won't know the difference between one that's stuck and won't boot and others. Plus many customers immediately make tickets instead of trying to power up the VPS after it goes offline so in any case having them powered on has more benefits than keeping them offline.
My VPS from NL ==> FFxxx ==> FF?? seems to be up and running after experiencing disk errors, and a long-ish downtime.
and a brand new YABS coming up just to make @cybertech happy.
Side benefits of being a @Virmach customer : one can develop a high threshold for patience (or patscience provider from Romania who-shall-not-be-named used to say).
And that's not a criticism!
Under US $ 10 a year is way cheaper than paying for meditation app/ Yog classes
Ban Reason: Banned for 34 login attempts.
Ban Expires: 07/08/2022 (21:00)
My multi-IP NY is down, likely moved and I wanted to find out what the score was. :'(
[EDIT1:]
One VPN later (ironically in NYC) and yep my triple IP VPS is fubar. Pointless changing main IP without the others being available. (NYCB018)
Awaiting non-patiently @VirMach
[EDIT2:]
At least the CHI is still there, for now.
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural. It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
@VirMach said: Before I change LAX to also have the fine-tuned NIC driver configuration does anyone who has both LAX and somewhere else functional notice one performing better than the other? On my end, LAX is around 70% cleaner.
Finally got my VPS on LAXA031 connected to the Internet, by clicking either fix networking or reinstallation buttons. Checking with "sar -n DEV 2 5", the result is far better (lower) on LAX (<100) than TYO(~1000 or more, even on the relatively stable node I previously mentioned). Nodes that behaving poorly like TYOC029 and TYOC002S could benefit a lot if the fix is applied on them, especially once you added the 10gbps switch to the storage node.
My vps1 on SJCZ006 had an automatic IP change from 213.59.116.x to 45.88.178.x.
The strange thing is, this happened when I typed sudo reboot in the OS.
Is VirBot giving out dynamic IP now, like residential network?
We accept Karma donations for the last flan. 🍮 affbrr
@VirMach said:
We're ramping up the abuse script temporarily. If you're at around full CPU usage on your VPS for more than 2 hours, it'll be powered down. We have too many VMs getting stuck on OS boot and negatively affecting others at this time, due to the operating systems getting stuck on boot for some after the Ryzen change. I apologize in advance for any false positives but I do want to note that it's technically within our terms for 2 hours of that type of usage to be considered abuse, we just usually try to be much more lenient.
Boot loop after migrating to a different CPU or changing to a different IP is not abuse.
Customer purchased service on a specific CPU and a specific IP that are not expected to change.
The kernel and userland could have been compiled with -march=native so that it would not start on any other CPU.
The services could have been configured to bind to a specific IP, which would cause service restart loop if the IP disappeared.
The safest way is not automatically powering on the service after the migration.
The customer needs to press Power On button themselves and then fixes the machine right away.
Running -march=native code on an unsupported CPU triggers undefined behavior.
Undefined behavior means anything could happen, such as pink unicorn appearing in VirMach offices, @deank stopping to believe in the end, or @FrankZ receiving 1000 free servers.
The simple act of automatic powering on a migrated server could cause these severe consequences and you don't want that.
We're ramping up the abuse script. It's what it is called. I didn't say boot loop after migrating is abuse.
Abuse script will just power it down, not suspend. I don't see the harm in powering down something stuck in a boot loop. I was just providing this as a PSA for anyone reading who might be doing something else not related that's also using a lot of CPU and for general transparency, we're making the abuse script more strict to try to power down the ones stuck in the boot loop automatically more quickly.
The safest way is not automatically powering on the service after the migration.
The customer needs to press Power On button themselves and then fixes the machine right away.
Not possible, we have to power up all of them to fix other issues. Otherwise we won't know the difference between one that's stuck and won't boot and others. Plus many customers immediately make tickets instead of trying to power up the VPS after it goes offline so in any case having them powered on has more benefits than keeping them offline.
Is FFME004 still down for everyone else? I notice it isn't mentioned in the latest status update. I get "The host is currently unavailable" in SolusVM, so I'm assuming it is more than me and haven't opened a ticket.
@tetech said:
Is FFME004 still down for everyone else? I notice it isn't mentioned in the latest status update. I get "The host is currently unavailable" in SolusVM, so I'm assuming it is more than me and haven't opened a ticket.
My VM in FFME004 is fine, but VMs in FFME005 & FFME006 are still getting the "no bootable device" error
@tetech said:
Is FFME004 still down for everyone else? I notice it isn't mentioned in the latest status update. I get "The host is currently unavailable" in SolusVM, so I'm assuming it is more than me and haven't opened a ticket.
My VM in FFME004 is fine, but VMs in FFME005 & FFME006 are still getting the "no bootable device" error
Networking Update - Just essentially waiting on DC hands at this point for NYC, Dallas, Atlanta, San Jose, Seattle, Phoenix, and Denver. Once these switch over networking should improve drastically and it should also have a positive impact on the CPU steal. We were supposed to have it done today but doesn't look likely they'll get to it, we'll see.
For Tokyo, we'll try to get it scheduled and completed by Tuesday. That one wasn't planned initially but we already cleaned up the IP blocks so we might as well move forward quickly.
Frankfurt AFAIK is already done that way. Frankfurt is having issues connecting to our servers and even connecting to me here in Los Angeles but networking looks superb on every other check so I'm guessing some common carrier pooched something up. This means it'll get a lot of panel errors unfortunately for the time being.
Disk Update - Frankfurt looks a lot better, I got the configurations to be mostly stable. Some Dallas, Los Angeles, and others also got Gen4 NVMe so as they have problems we're already on it and fixing them but feel free to report any issues for the disks. I/O related errors only please, not being offline or never going online, we already have that part fully figured out and in the works as well, on a higher level.
Tokyo Storage Update - Haven't been able to get to it.
NYC Storage Update - Working on fast-tracking this as it's been heavily delayed and we really need it. We've already moved backups away from the same company the servers are with just in case. They're definitely not taking it well that we're actually leaving, like a crazy ex. It's potentially turning into a Psychz nightmare scenario but that's pretty much all I can say on that. So rest assured if they do pull anything we've got it covered. Of course it's not perfect. They're also being very nosey.
Amsterdam Storage Update - Same as above but secondary level of urgency meaning we want to get out NYC Metro first. I'll make it up to you guys for waiting so long, just remind me if I don't.
Template Syncs - Ongoing, many more re-synced. OS installs should work better. QuadraNet's "DDoS Protection" is essentially just hefty false positives though so for Los Angeles I might have to literally drive down a hard drive and load them on at this point since all the tweaks they do still don't allow for it. This also got in the way of a lot of backups and huge headaches with migrations.
Windows template is fully broken at this point, I think I synced a bad version. Looking into that later today hopefully.
In amongst all this chaos, no further updates on rDNS nor multi-IP, @Virmach ?
Just askin'.
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural. It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
Re: Template Syncs, I'll give my 256MB fubar'ed Dallas another attempt at a reinstall. (Not that this VPS is important to me, now.)
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural. It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
@AlwaysSkint said:
In amongst all this chaos, no further updates on rDNS nor multi-IP, @Virmach ?
Just askin'.
Bottom of the list at this point to be frank. I understand it's very important to some, we just have to make sure people have a functional service first, a functional IP second, that the networking actually functions well, more builds, coordinating shipments, backups, migrations, and getting through the literal thousands of tickets and damage control.
It'll probably go all of the above, then IPv6, then rDNS, then multi-IP. Likely sometime in August and I'm trying to be conservative but you know how that goes. Honestly though if we can't get it done by the end of August and you need multiple IP or rDNS... I'd be totally on your side if you were furious.
Side note on multi-IP, at this point it's going to require a lot of work to sort through all of it and that's after we set up more nodes. Multi-IP will most likely for most people require another migration as well. Originally this wasn't going to be a problem but originally we were also counting on IP subnet overlaps. Now that we're locking everything down to a subnet per node, it becomes tremendously more difficult.
Heart wouldn't like that, so more likely very frustrated. [Finally managed to walk 2 miles today - yipee!]
Bounced server cron/system emails are a nuisance due to lack of rDNS. The lack of multi-IP would seal the fate of that one particular VPS of mine, as it would no longer be cost effective, so to speak.
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural. It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
..Still get boot failure from CD on DALZ007, when trying a Ryzen Debian 10 (and likely others).
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural. It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
@VirMach said: Template Syncs - Ongoing, many more re-synced. OS installs should work better. QuadraNet's "DDoS Protection" is essentially just hefty false positives though so for Los Angeles I might have to literally drive down a hard drive and load them on at this point since all the tweaks they do still don't allow for it.
I'm sure this got dropped in all the billion other things going on but since you're talking about template syncs I figured I'd bring it back up. BF-SPECIAL-2020 only shows the stock four ( C7, C8, Deb8, Deb9 ) for mountable ISOs. Absolutely not a show-stopper or anything just wanted to make sure it was on a list somewhere.
I have an older SJC VM that rebooted this morning, stayed up for a while, but has been down most of the day with billing panel saying "the node is currently locked". This VPS is not on a Ryzen node unless it has been migrated (I haven't kept track but I didn't request a migration, figuring I'd wait til the smoke clears). "Server information" only says that the node is locked, and doesn't give the other info such as the node name, so I don't know what node it is on.
It also looks like my Ryzen vps has gotten rebuilt or something like that. I don't mean the rebuild from April but something more recent. Did that happen too? Again I haven't followed discussion that closely. This one is on SJCZ005.
Is the non-Ryzen stuff already known to be having issues? It has been working ok up til today.
@willie said:
1. I have an older SJC VM that rebooted this morning, stayed up for a while, but has been down most of the day with billing panel saying "the node is currently locked". This VPS is not on a Ryzen node unless it has been migrated (I haven't kept track but I didn't request a migration, figuring I'd wait til the smoke clears). "Server information" only says that the node is locked, and doesn't give the other info such as the node name, so I don't know what node it is on.
It also looks like my Ryzen vps has gotten rebuilt or something like that. I don't mean the rebuild from April but something more recent. Did that happen too? Again I haven't followed discussion that closely. This one is on SJCZ005.
Is the non-Ryzen stuff already known to be having issues? It has been working ok up til today.
This was supposed to be a quick migration but QuadraNet has been nullrouting our IP addresses all day for doing the transfers. I guess their DDoS protection is some script they set up to deny you service, the 14 year old developer must have misheard the objective.
I've been going through these more closely and they nullrouted someone for 5 days for sending 60MB/s traffic to their backup VPS for 10 minutes. Every single migration today has been completely botched by that, I have to go through them again and make sure it didn't also corrupt the data it transferred.
It's also possible you're part of the other migration which is going slow for other reasons. CC San Jose has always had weird routing problems and on top of that the Ryzens are on a quarter cabinet that are especially not doing well with the bigger VLANs right now. I assume they're using some old half dead switch (for quarter cabs they provide ports.) Maybe I'm just upset as a result of sleep deprivation, but if I had any energy left this year I'd say it looks like we need to do another round of migrations again already. But hopefully they'll have it resolved soon or else we'll have to start doing migrations via UPS overnighted USB drives to improve efficiency.
@willie said: I have an older SJC VM that rebooted this morning, stayed up for a while, but has been down most of the day with billing panel saying "the node is currently locked".
I expect your VM is being migrated to Ryzen currently. The normal billing panel will show up again after the migration is complete.
Comments
Just in case you're thinking this is about the migrations, the discussion you're referring back to is actually about Ryzen Location Change button and @yoursunny opinions on it.
Ok, I understand that, but is there a way to fix network manually? What settings should i set to get vm vorking? I couldn't see any problems from the first approach - there is eth0 ( ens3) interface up, there is a route to gateway xxx.xxx.xxx.1
I can ping nearby nodes, but for gateway response is Network host unreachable.
I understand, I'm just adding onto it that I agree with you and it's not user friendly and I'm also saying that unfortunately our current situation is similar and also not how we intended for it to be coded by the developer. I did read what @yoursunny suggested and that definitely has its benefits but our ideal version would be immediate.
I remember we improved our script to minimize downtime to only something like an hour per VM and had that concept working pretty well, but it didn't work out when you're planning for efficiency and many servers.
How the project was for the developer that never completed it:
It'd have been pretty nice if we had that ready in time for these migrations so it instead didn't end up as the vague day-long periods.
Trying to figure out from network status and backlog here: Should FFME002 be operational?

Have never gotten it to work here ...
Troubleshooter reports:
Main IP pings: false
Node Online: false
Service online: offline
Operating System: linux-ubuntu-16.04-server-x86_64-minimal-latest
Service Status:Active
pokes FFME002 with a stick nope, still dead. has been for a while, worked for like most of a day after i migrated there then dead
I see new Los Angeles Ryzen network better... 10k in traffic vs 100k in Atlanta
Something strange about FFME003 network config of my vm. If i set network interface to dhcp, i receive correct ip address (xxx.xxx.163.xxx) but from dhcp server of another subnet xxx.xxx.162.2. And ip address i receive is from the same subnet as FFME004 vm. Is this working as intended? If i restore network config from control panel, i get the same ip as static, excluding missing symlink to resolvconf.
Most services are paid annually.
One migration per year is not enough.
That's why I suggested once per month.
Some starter credits should be granted on the current services, because:
Once all the looking glass nodes are up, can we at least have 2~3 credits per year?
That's a lot more flexible than only one credit per year, in case the network condition deteriorates in the middle.
We accept Karma donations for the last flan. 🍮 affbrr
Seems many JP VMs got their IP changed without prior notice
Indeed. No email, and no information I could find on the Network Status page.
While I understand that a host can face many challenges that cannot be predicted or immediately explained, this is not one of such situations. And the lack of a warning and the lack of information about why the change happened (Was it intentional and is it to stay, or was was it just a configuration mistake that will be reverted?) is a big fat NO NO in my book.
What tha F is going in with FFME? Migrated from FFME004 yesterday for 3 bucks fee, got the same FFME004, but online and working, today it's totally offline - no boot, no VNC, nothing. And no information about what's going on. I was waiting patiently for two weeks, created zero tickets, but now i have to give up and leave as soon as i could get my data from FFME003. Honestly, even my home server with my hobbyist approach does have less downtime and more reliability, and migrates faster.
While our Control Panel still shows 0 IPv4 addresses, we can finally get back in at the assigned IPv4 address and also through our private VPN. So progress has been made at least on the RYZE.SEA-Z002.VMS node in Seattle. We'll see if it stays up.
We're ramping up the abuse script temporarily. If you're at around full CPU usage on your VPS for more than 2 hours, it'll be powered down. We have too many VMs getting stuck on OS boot and negatively affecting others at this time, due to the operating systems getting stuck on boot for some after the Ryzen change. I apologize in advance for any false positives but I do want to note that it's technically within our terms for 2 hours of that type of usage to be considered abuse, we just usually try to be much more lenient.
We did send out emails, but it's possible they did not all send due to SolusVM also being overloaded around that time. Please also check your spam box, we sent these directly from SolusVM.
Well initially the way it's going to work is there will be a period of time where you're allowed to "Ryzen Migrate" to your desired location (without data) as everyone lands in their desired location. It'll be announced here and on OGF, as well as most likely an "Announcement" on our website and probably a 1-2 week period where this can be done by everyone eligible.
The credits I described are for a system not yet coded.
I'll make a network status page for it since it seems a lot of the emails failed to send.
FFME004, we found ECC error, and the settings also dropped off again. Memory swap fixed FFME004, we couldn't send out migration emails in time because xTom worked very quickly to get this replaced. The setting drop-off caused a disk to drop and node was online but VMs were not booting. That has been resolved. I'm checking it again to see if settings stick, if they don't there might be another reboot which we'll create network status for but hopefully this will be stable moving forward.
FFME has had 3 network status, many updates here, on OGF, and probably a fair share of emails. It can be considered an ongoing issue until they prove themselves by remaining online for more than 2 days.
Settings stuck on FFME004 but I'm pretty sure I've said that once before. There's zero information on this but there's been constant kernel bugs regarding these fixes on Linux. I can't rewrite the Linux kernel right now so until Linux figures out how it's going to treat these problems I don't know what else to do about it.
These issues have come on gone ever since NVMe SSDs existed, you can search online since around 2015. If you just look at Linux bug trackers it seems like every version fixes one thing and breaks another. I'll have to come up with some kind of kernel update plan moving forward where we try to mitigate the issues from re-appearing. But of course the solution can't just be to stay on the same version for 5 years.
We're using literal copies of the same nodes in FFM... in Tokyo, and they're not having the same problems with the only difference being kernel versions.
Spoke too soon. Status back to dead in the water this afternoon.
My VM in FFME004 seems fine now, but I have a couple of VMs FFME005 and FFME006 having Status "Offline", can't bootup, can't reinstall OS. Hope you can help, thanks! Ticket #754039
Boot loop after migrating to a different CPU or changing to a different IP is not abuse.
Customer purchased service on a specific CPU and a specific IP that are not expected to change.
The kernel and userland could have been compiled with
-march=nativeso that it would not start on any other CPU.The services could have been configured to bind to a specific IP, which would cause service restart loop if the IP disappeared.
The safest way is not automatically powering on the service after the migration.
The customer needs to press Power On button themselves and then fixes the machine right away.
Running
-march=nativecode on an unsupported CPU triggers undefined behavior.Undefined behavior means anything could happen, such as pink unicorn appearing in VirMach offices, @deank stopping to believe in the end, or @FrankZ receiving 1000 free servers.
The simple act of automatic powering on a migrated server could cause these severe consequences and you don't want that.
We accept Karma donations for the last flan. 🍮 affbrr
This has an issue with a software getting stuck and duplicating its process over and over until it overloads and we have to reboot it. We made some changes, if it happens again we'll try to catch it earlier this time to avoid a reboot.
Were these always offline after Ryzen Migrate button?
We're ramping up the abuse script. It's what it is called. I didn't say boot loop after migrating is abuse.
Abuse script will just power it down, not suspend. I don't see the harm in powering down something stuck in a boot loop. I was just providing this as a PSA for anyone reading who might be doing something else not related that's also using a lot of CPU and for general transparency, we're making the abuse script more strict to try to power down the ones stuck in the boot loop automatically more quickly.
Not possible, we have to power up all of them to fix other issues. Otherwise we won't know the difference between one that's stuck and won't boot and others. Plus many customers immediately make tickets instead of trying to power up the VPS after it goes offline so in any case having them powered on has more benefits than keeping them offline.
The "Migration" button was not used. Has been always offline after the migration. It happened after the planned migration from AMS to FFE
My VPS from NL ==> FFxxx ==> FF?? seems to be up and running after experiencing disk errors, and a long-ish downtime.
Side benefits of being a @Virmach customer : one can develop a high threshold for patience (or patscience provider from Romania who-shall-not-be-named used to say).
And that's not a criticism!
Under US $ 10 a year is way cheaper than paying for meditation app/ Yog classes
Cheers
blog archives
Fantastic!
My multi-IP NY is down, likely moved and I wanted to find out what the score was. :'(
[EDIT1:]
One VPN later (ironically in NYC) and yep my triple IP VPS is fubar. Pointless changing main IP without the others being available. (NYCB018)
Awaiting non-patiently @VirMach
[EDIT2:]
At least the CHI is still there, for now.
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural.
It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
IIRC my FFME0002 node was online for a while after hitting the migrate button. But then it's been dead ever since ...

Yep, same. I poke mine with a stick occasionally but it hasn't really done anything of note since the first day.
Finally got my VPS on LAXA031 connected to the Internet, by clicking either fix networking or reinstallation buttons. Checking with "sar -n DEV 2 5", the result is far better (lower) on LAX (<100) than TYO(~1000 or more, even on the relatively stable node I previously mentioned). Nodes that behaving poorly like TYOC029 and TYOC002S could benefit a lot if the fix is applied on them, especially once you added the 10gbps switch to the storage node.
Yabs on LA node LAXA031, looks good:
Mine is up now!

Yep, I reinstalled mine from scratch (that one is a tiny bf so not much on it) and it's still up now hours later. crosses fingers
My vps1 on SJCZ006 had an automatic IP change from 213.59.116.x to 45.88.178.x.
The strange thing is, this happened when I typed
sudo rebootin the OS.Is VirBot giving out dynamic IP now, like residential network?
We accept Karma donations for the last flan. 🍮 affbrr
Please look ticket #618039
Synteq Technical Support, Technical Writer.
Contact me at: +1 (307) 428 8111 or [email protected]
Is FFME004 still down for everyone else? I notice it isn't mentioned in the latest status update. I get "The host is currently unavailable" in SolusVM, so I'm assuming it is more than me and haven't opened a ticket.
My VM in FFME004 is fine, but VMs in FFME005 & FFME006 are still getting the "no bootable device" error
Oh, interesting. Maybe it is time for a ticket.
Networking Update - Just essentially waiting on DC hands at this point for NYC, Dallas, Atlanta, San Jose, Seattle, Phoenix, and Denver. Once these switch over networking should improve drastically and it should also have a positive impact on the CPU steal. We were supposed to have it done today but doesn't look likely they'll get to it, we'll see.
For Tokyo, we'll try to get it scheduled and completed by Tuesday. That one wasn't planned initially but we already cleaned up the IP blocks so we might as well move forward quickly.
Frankfurt AFAIK is already done that way. Frankfurt is having issues connecting to our servers and even connecting to me here in Los Angeles but networking looks superb on every other check so I'm guessing some common carrier pooched something up. This means it'll get a lot of panel errors unfortunately for the time being.
Disk Update - Frankfurt looks a lot better, I got the configurations to be mostly stable. Some Dallas, Los Angeles, and others also got Gen4 NVMe so as they have problems we're already on it and fixing them but feel free to report any issues for the disks. I/O related errors only please, not being offline or never going online, we already have that part fully figured out and in the works as well, on a higher level.
Tokyo Storage Update - Haven't been able to get to it.
NYC Storage Update - Working on fast-tracking this as it's been heavily delayed and we really need it. We've already moved backups away from the same company the servers are with just in case. They're definitely not taking it well that we're actually leaving, like a crazy ex. It's potentially turning into a Psychz nightmare scenario but that's pretty much all I can say on that. So rest assured if they do pull anything we've got it covered. Of course it's not perfect. They're also being very nosey.
Amsterdam Storage Update - Same as above but secondary level of urgency meaning we want to get out NYC Metro first. I'll make it up to you guys for waiting so long, just remind me if I don't.
Template Syncs - Ongoing, many more re-synced. OS installs should work better. QuadraNet's "DDoS Protection" is essentially just hefty false positives though so for Los Angeles I might have to literally drive down a hard drive and load them on at this point since all the tweaks they do still don't allow for it. This also got in the way of a lot of backups and huge headaches with migrations.
Windows template is fully broken at this point, I think I synced a bad version. Looking into that later today hopefully.
In amongst all this chaos, no further updates on rDNS nor multi-IP, @Virmach ?
Just askin'.
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural.
It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
Re: Template Syncs, I'll give my 256MB fubar'ed Dallas another attempt at a reinstall. (Not that this VPS is important to me, now.)
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural.
It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
Bottom of the list at this point to be frank. I understand it's very important to some, we just have to make sure people have a functional service first, a functional IP second, that the networking actually functions well, more builds, coordinating shipments, backups, migrations, and getting through the literal thousands of tickets and damage control.
It'll probably go all of the above, then IPv6, then rDNS, then multi-IP. Likely sometime in August and I'm trying to be conservative but you know how that goes. Honestly though if we can't get it done by the end of August and you need multiple IP or rDNS... I'd be totally on your side if you were furious.
Side note on multi-IP, at this point it's going to require a lot of work to sort through all of it and that's after we set up more nodes. Multi-IP will most likely for most people require another migration as well. Originally this wasn't going to be a problem but originally we were also counting on IP subnet overlaps. Now that we're locking everything down to a subnet per node, it becomes tremendously more difficult.
Heart wouldn't like that, so more likely very frustrated.
[Finally managed to walk 2 miles today - yipee!]
Bounced server cron/system emails are a nuisance due to lack of rDNS. The lack of multi-IP would seal the fate of that one particular VPS of mine, as it would no longer be cost effective, so to speak.
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural.
It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
..Still get boot failure from CD on DALZ007, when trying a Ryzen Debian 10 (and likely others).
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural.
It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
Miami beach club update - DC hands are being distracted by sexy FrankZ dancing on the bus.
We accept Karma donations for the last flan. 🍮 affbrr
I'm sure this got dropped in all the billion other things going on but since you're talking about template syncs I figured I'd bring it back up. BF-SPECIAL-2020 only shows the stock four ( C7, C8, Deb8, Deb9 ) for mountable ISOs. Absolutely not a show-stopper or anything just wanted to make sure it was on a list somewhere.
I have an older SJC VM that rebooted this morning, stayed up for a while, but has been down most of the day with billing panel saying "the node is currently locked". This VPS is not on a Ryzen node unless it has been migrated (I haven't kept track but I didn't request a migration, figuring I'd wait til the smoke clears). "Server information" only says that the node is locked, and doesn't give the other info such as the node name, so I don't know what node it is on.
It also looks like my Ryzen vps has gotten rebuilt or something like that. I don't mean the rebuild from April but something more recent. Did that happen too? Again I haven't followed discussion that closely. This one is on SJCZ005.
Is the non-Ryzen stuff already known to be having issues? It has been working ok up til today.
This was supposed to be a quick migration but QuadraNet has been nullrouting our IP addresses all day for doing the transfers. I guess their DDoS protection is some script they set up to deny you service, the 14 year old developer must have misheard the objective.
I've been going through these more closely and they nullrouted someone for 5 days for sending 60MB/s traffic to their backup VPS for 10 minutes. Every single migration today has been completely botched by that, I have to go through them again and make sure it didn't also corrupt the data it transferred.
It's also possible you're part of the other migration which is going slow for other reasons. CC San Jose has always had weird routing problems and on top of that the Ryzens are on a quarter cabinet that are especially not doing well with the bigger VLANs right now. I assume they're using some old half dead switch (for quarter cabs they provide ports.) Maybe I'm just upset as a result of sleep deprivation, but if I had any energy left this year I'd say it looks like we need to do another round of migrations again already. But hopefully they'll have it resolved soon or else we'll have to start doing migrations via UPS overnighted USB drives to improve efficiency.
I expect your VM is being migrated to Ryzen currently. The normal billing panel will show up again after the migration is complete.