Okay after working on this all day, outside of the two nodes somehow still down (SJCZ004 and NYCB036) I've counted these specifically and we've gone down from something like 2% of people after the most recent round of migrations to exactly 0.91% of people on Ryzen having a VPS with an incorrect IP or non-bootable. At this point for the other two I basically just need the DC hands to reboot them at least and then I'll probably be able to try to take it from there... I hope it doesn't end up taking two weeks for that type of request but we'll see as we're on our way there.
All the others left have been organized and require a re-migration at this point. This number also does include all the broken VMs from Ryzen migrate button, which make up the majority and are clustered on 28 different nodes. What we're going to do is first create the LVMs manually to fix the majority and boot them up, and then either within 24 hours try to restore the data if it exists, or if it takes any longer, send out a ticket and ask people if they still wanted their old data (in case they begin using it and already loading in their own or just don't have important data.)
This last bit took astronomically longer to identify and fix but now that it's the ending portion of it and heavily organized with all the other problems (the other 1% fixed today) out of the way, it should go smoothly.
Then we'll probably run our auto credit and close ticket script again for everyone, probably piss off a few hundred people at least, but actually be able to get on track with tickets as otherwise the majority of the tickets are going to end up being for issues we already resolved. I'll try to wait on this until we get the other two nodes back up and finalized as well and any remaining migrations done.
From the fact that I still have a number of vps that haven't migrated yet, I'm guessing there is a certain amount total that haven't changed yet. I'm assuming the future migrations left hopefully will be to existing systems that have been tested?
@Daevien said:
From the fact that I still have a number of vps that haven't migrated yet, I'm guessing there is a certain amount total that haven't changed yet. I'm assuming the future migrations left hopefully will be to existing systems that have been tested?
All systems are existing systems that have been tested, some more than others. The newer nodes I've actually spent less time testing, not as a result of being careless, but in that a lot of the issues were already ironed out so it was a quicker fix and it's possible less of these core issues get passed onto those nodes. We have not sold anything on these Ryzen nodes for several weeks now so a good portion of migrations will be to existing nodes that freed up a little bit. The rest will mostly be Hivelocity setups and the final bit will be the last servers I've built which essentially have the best solid state drives (IMO) so they shouldn't "drop off" and the motherboards have been more extensively tested and pre-flashed (mostly.) Any issues these nodes may have will be related to power and brackets falling apart during shipping which get fixed before it goes online. I'm being specific about where I deploy them to ensure if we run into problems, it should hopefully be a quick fix. Only caveat is that the Hivelocity locations will have zero RAM available. We had to do so many RAM swaps and they're all with existing partners, I ran out, but I memory tested everything.
Networking configuration will be the one that gives us no problems straight out the gate for the new nodes as well and I've not sent any additional servers outside of storage to the partner who shall not be named who has taken over a week for a power button press and two weeks for switch configuration or initial node deployments.
We've also gotten quite good at migrations by now so those will hopefully be as smooth as they get for the last bit.
I did also finally get some switch access, so you will either see more nodes up in NYC that we can use for migrations, or you'll see me break the networking.
Is PHXZ004 working fine? I couldn't log in via SSH, so I logged in via VNC, but even ping to xxx.xxx.xxx.1 failed.
No reply required. Instead of writing a reply to me, you can rest your eyes and body 😉
I was going to get a flight to San Jose for tomorrow morning but I'm still waiting to see how we can even get DC access since it's the first time. I might still go on Tuesday and see if I can beat the DC hands to it for SJCZ004 at this point, only about an hour flight.
@tototo said:
Is PHXZ004 working fine? I couldn't log in via SSH, so I logged in via VNC, but even ping to xxx.xxx.xxx.1 failed.
No reply required. Instead of writing a reply to me, you can rest your eyes and body 😉
Phoenix actually has some of our infrastructure on it and I can confirm it's being terrible. But it's still mostly usable. We need to redo network configuration on it still.
all my VPS (LA) can't ping to 8.8.8.8. any explanation on that?
Very unlucky user or very lucky user based on preference, when it happened, etc. Maybe your services are getting migrated to Ryzen, maybe they're all broken.
NYC storage node, I had to do a lot of weird setups for this one to make it work as we faced problems such as a disk going missing, the wrong switch, and so on in the background. Right now it's LACP aggregated 3Gbps.
Originally was supposed to have 36 disks, then had 31, then Amazon delayed it, then either one died or DC lost it so we had 30 and I can't keep waiting at this point to try to do 36 disk RAID. This is a beefier card, I figured we give it a try with two 16 disk arrays to essentially emulate two Tokyo smaller servers or to add further fun challenges if I'm wrong about this decision. I configured dm-cache but I don't know what to do with it yet outside of just maybe using it for a test VM. SolusVM surprisingly added this feature in V2, whenever that ends up happening we'll be ready and you can get a whopping 7GB of Gen4 NVMe cache with your 1TB storage (I can probably get it it up higher it has open slots.) Don't know if it's any good.
This is a Threadripper 3960X. It's the infamous motherboard I dropped, I killed two of the RAM slots so that's why it's at 192GB instead of the planned 256GB and why 36 disk --> 2 x 16 (32) makes more sense, but I ran a lot of tests on it and it was otherwise healthy.
This 9560 has a weird battery (CacheVault) non-issue, I think it just takes longer to boot or has a more strict requirement before initiating. It's been like that since brand new and I've looked into it in every way I could with the limited time, but it only happens on initial boot and shows healthy state.
Fun fact though, it's now impossible to get a replacement motherboard for these, fully. I've looked so hard. So if it ever dies we'll just do a board swap with Epyc, 5950X, 6950X by then, or something like that.
9560-16i RAID controller.
Oh and we already have like 3 NICs now to make it 10Gbps, 20Gbps and later 40Gbps even maybe. We basically spam ordered and crossed fingers for this one since it's initially a lot larger than Tokyo and needs the 10Gbps. They just need to connect it.
I'm about to send the maintenance email for one of the current storage nodes, while we were running backups this week two of the disks on it died and it's degraded. That's one reason I was rushing this out in the last 2 days, I am absolutely not getting DC hands involved at the old DC, every single RAID-related thing they've done recently has been guaranteed destroying all data and the last time they worked on this they almost jumbled the data. Luckily that means we already have like 80% of a massive storage server already backed up to another storage server and I'll be moving off the people that didn't get backed up first.
Alright now let's fill it and bring those numbers down. I'll also try to start moving the other one less aggressively to LAX storage and shifting our disaster recovery backups from that to non-RAID since they're mostly complete. Once we finish the rest of the NY storage we'll let you move back and I apologize in advance for the double move but it's better than it getting yoinked offline.
What is the consensus with networking in Phoenix (PHXZ003), are whole nodes having issues? I cannot ping to 8.8.8.8 from the VPS after trying the usual stuff like a hard power-off then on. I have the "Fix Ryzen IP" option, but don't want to just try stuff randomly if the whole site is having issues.
@tetech said:
What is the consensus with networking in Phoenix (PHXZ003), are whole nodes having issues? I cannot ping to 8.8.8.8 from the VPS after trying the usual stuff like a hard power-off then on. I have the "Fix Ryzen IP" option, but don't want to just try stuff randomly if the whole site is having issues.
my LAX had 8.8.8.8 issues for a while , but went back up.
funny enough i was trying to do a bench.sh test and it was coincidentally down, and up afterwards.
@tetech said: What is the consensus with networking in Phoenix (PHXZ003), are whole nodes having issues? I cannot ping to 8.8.8.8 from the VPS after trying the usual stuff like a hard power-off then on. I have the "Fix Ryzen IP" option, but don't want to just try stuff randomly if the whole site is having issues.
Not just you, I'm on PHXZ002 and the network was so so before, but now no network at all for last ~24 hours. I don't think you can fix this yourself. VirMach mentioned issues in Phoenix above.
@VirMach said: Phoenix actually has some of our infrastructure on it and I can confirm it's being terrible. But it's still mostly usable. We need to redo network configuration on it still.
@willie said:
Looks like my vpsshared site is down, not responding to pings. It's a super low traffic personal site, so I'll survive, but it sounds like things are probably backlogged there.
I'd scream at CC but I've given up on that. They placed a permanent nullroute on the main IP. This is probably the 30th time they've done this, from a single website being malicious. I'm trying to get a server up and just move it at this point.
Again I'd usually be freaking out and working on it immediately to get it online but I'm just being realistic here when there's probably a dozen others in worse states. It sucks, it's unprofessional, but there's no use having more stress over it as that'll just slow me down.
I'm wondering if its LA10GKVM14 since that node has been down almost one month already
@tetech said:
What is the consensus with networking in Phoenix (PHXZ003), are whole nodes having issues? I cannot ping to 8.8.8.8 from the VPS after trying the usual stuff like a hard power-off then on. I have the "Fix Ryzen IP" option, but don't want to just try stuff randomly if the whole site is having issues.
my LAX had 8.8.8.8 issues for a while , but went back up.
funny enough i was trying to do a bench.sh test and it was coincidentally down, and up afterwards.
Honestly at this point I think 8.8.8.8 issue is actually not on any specific datacenter's end. We had it happen on our WHMCS for a little bit, and that one's not even hosted with any datacenter we use right now for these Ryzens. We've also had it reported on LAX which is QN and Phoenix which is PhoenixNAP IIRC.
So it has to be something upstream probably unless it's in some weird way related?
@willie said:
Looks like my vpsshared site is down, not responding to pings. It's a super low traffic personal site, so I'll survive, but it sounds like things are probably backlogged there.
I'd scream at CC but I've given up on that. They placed a permanent nullroute on the main IP. This is probably the 30th time they've done this, from a single website being malicious. I'm trying to get a server up and just move it at this point.
Again I'd usually be freaking out and working on it immediately to get it online but I'm just being realistic here when there's probably a dozen others in worse states. It sucks, it's unprofessional, but there's no use having more stress over it as that'll just slow me down.
I'm wondering if its LA10GKVM14 since that node has been down almost one month already
LA10GKVM14 was part of a massive PDU or power circuit event with them, where like a dozen plus servers went offline and online, and then offline, causing a bunch of power supplies to fail. A lot of times when we've had power issues like that, CC also moves them to another switch without telling us and then does not configure the networking properly. This is one of those cases, so it's basically been left without functional networking and no one's willing to help. We're marking it at a loss at this point because it's impossible to get them to do anything anymore. It's possible this also has other issues like a failed controller/data corruption as a result of it taking so long for them to hook up another PSU, I think it took 4 or 5 days and the battery drained for the cache but it doesn't matter. We have to try to locate backups which we haven't had luck with so far because backups also failed and were closely tied to the power event, I think we have partial backups.
So right now we're mostly stuck on regenerating services.
@tetech said: What is the consensus with networking in Phoenix (PHXZ003), are whole nodes having issues? I cannot ping to 8.8.8.8 from the VPS after trying the usual stuff like a hard power-off then on. I have the "Fix Ryzen IP" option, but don't want to just try stuff randomly if the whole site is having issues.
Not just you, I'm on PHXZ002 and the network was so so before, but now no network at all for last ~24 hours. I don't think you can fix this yourself. VirMach mentioned issues in Phoenix above.
@VirMach said: Phoenix actually has some of our infrastructure on it and I can confirm it's being terrible. But it's still mostly usable. We need to redo network configuration on it still.
@tetech said:
What is the consensus with networking in Phoenix (PHXZ003), are whole nodes having issues? I cannot ping to 8.8.8.8 from the VPS after trying the usual stuff like a hard power-off then on. I have the "Fix Ryzen IP" option, but don't want to just try stuff randomly if the whole site is having issues.
my LAX had 8.8.8.8 issues for a while , but went back up.
funny enough i was trying to do a bench.sh test and it was coincidentally down, and up afterwards.
Honestly at this point I think 8.8.8.8 issue is actually not on any specific datacenter's end. We had it happen on our WHMCS for a little bit, and that one's not even hosted with any datacenter we use right now for these Ryzens. We've also had it reported on LAX which is QN and Phoenix which is PhoenixNAP IIRC.
So it has to be something upstream probably unless it's in some weird way related?
i think so too. something related geographically in in US
@VirMach said: We have to try to locate backups which we haven't had luck with so far because backups also failed and were closely tied to the power event, I think we have partial backups.
So right now we're mostly stuck on regenerating services.
I think most of us/customers would be happy with fresh idlers servers. Specially after a month of downtime so I wouldn't spend any extra time for hunting backups but I usually to keep my own backups.
Tokyo is full right now, these will get activated by Wednesday most likely. If you'd like any of you guys (mucstudio, nauthnael, TrueBlumdfeld, gkl1368) can reply back to me here and request a refund or I can activate you in San Jose for now and migrate you to Tokyo when it's ready. Or of course you can just wait.
Could you please migrate me to Tokyo?
It's been offline for ten days on SJZ004 and the control pannel page timeout
Tokyo is full right now, these will get activated by Wednesday most likely. If you'd like any of you guys (mucstudio, nauthnael, TrueBlumdfeld, gkl1368) can reply back to me here and request a refund or I can activate you in San Jose for now and migrate you to Tokyo when it's ready. Or of course you can just wait.
Could you please migrate me to Tokyo?
It's been offline for ten days on SJZ004 and the control pannel page timeout
order is: 4678913601
How do you propose a migration of an offline server?
@VirMach said:
How do you propose a migration of an offline server?
Can you migrate my broken VPS in Phoenix to getRandomLocation() without data?
(Yes, needless to say, this is a joke. I'd rather wait than open a ticket)
Tokyo is full right now, these will get activated by Wednesday most likely. If you'd like any of you guys (mucstudio, nauthnael, TrueBlumdfeld, gkl1368) can reply back to me here and request a refund or I can activate you in San Jose for now and migrate you to Tokyo when it's ready. Or of course you can just wait.
Could you please migrate me to Tokyo?
It's been offline for ten days on SJZ004 and the control pannel page timeout
order is: 4678913601
How do you propose a migration of an offline server?
@VirMach said:
How do you propose a migration of an offline server?
Can you migrate my broken VPS in Phoenix to getRandomLocation() without data?
(Yes, needless to say, this is a joke. I'd rather wait than open a ticket)
I'd actually accept these kinds of request if it was possible to keep it clean but realistically that means you get a new VM and then the old one just kind of hangs around while it's offline and takes up space until we manually verify which ones are abandoned and clear up the space.
For Phoenix though it's technically online so we could start allowing Ryzen migrations. Let me just finish up with what I need to do this morning and I'll activate the button again.
@VirMach said:
How do you propose a migration of an offline server?
Can you migrate my broken VPS in Phoenix to getRandomLocation() without data?
(Yes, needless to say, this is a joke. I'd rather wait than open a ticket)
I'd actually accept these kinds of request if it was possible to keep it clean but realistically that means you get a new VM and then the old one just kind of hangs around while it's offline and takes up space until we manually verify which ones are abandoned and clear up the space.
For Phoenix though it's technically online so we could start allowing Ryzen migrations. Let me just finish up with what I need to do this morning and I'll activate the button again.
I think it would help some users, but "(non-Phoenix location) has no buttons!!!" and that could open a ticket.
I am not really in a hurry and can wait. That said, I respect your decision.
@Virmach
Additional 7 IPs to VPS two months ago. Some additional IPs have been unavailable for about a month. The additional IPs of the billing panel also cannot be displayed. I opened a ticket on July 6th. I'd appreciate it if you could help me check it out.
Ticket #218767
@VirMach said:
How do you propose a migration of an offline server?
Can you migrate my broken VPS in Phoenix to getRandomLocation() without data?
(Yes, needless to say, this is a joke. I'd rather wait than open a ticket)
I'd actually accept these kinds of request if it was possible to keep it clean but realistically that means you get a new VM and then the old one just kind of hangs around while it's offline and takes up space until we manually verify which ones are abandoned and clear up the space.
For Phoenix though it's technically online so we could start allowing Ryzen migrations. Let me just finish up with what I need to do this morning and I'll activate the button again.
I think it would help some users, but "(non-Phoenix location) has no buttons!!!" and that could open a ticket.
I am not really in a hurry and can wait. That said, I respect your decision.
Button updated.
I can't get it to work but let's see if anyone else has any luck (edit -- looks like it did work for at least one person so far.) To be far I was trying to break it by being impatient and refreshing/closing it. Let me allow it to run for 5 minutes and see what happens instead.
Oh wait I think it's because I activated our trap card
Would it be possible to show that button also ppl with servers in broken nodes (like LA10GKVM14)? That way users could regenerate their services in different nodes
@atomi said:
Would it be possible to show that button also ppl with servers in broken nodes (like LA10GKVM14)? That way users could regenerate their services in different nodes
There's no way to differentiate between fully broken and partially broken nodes to display it only to those people.
@VirMach said: I'd actually accept these kinds of request if it was possible to keep it clean but realistically that means you get a new VM and then the old one just kind of hangs around while it's offline and takes up space until we manually verify which ones are abandoned and clear up the space.
I actually migrated out of a broken Chicago node. I have two machines as a result just one is unavailable.
@FrankZ said:
Did I miss the Ryzen migrate button for Phoenix or was that not a thing?
EDIT: Never mind Phoenix network seems to be working now, slow, but working.
Took it away while we do more migrations, we can't have it drastically change too much or it'll put the plans out of whack. It'll re-enable later tonight.
Finally got back the physical port numbers for the servers I needed, NYCB036 (formerly 101) should be fixed soon, just need to wrap up what I was working on and configure it in the switch. SJCZ004 unfortunately still down, I was going to fly in yesterday and just literally press the button then fly back but flight times didn't work out with the required 24 hour notice to access the site.
DALZ009 appears to be having problems since at least last night, didn't get a chance to send out emails or look into it yet. LAXA014 still keeps rebooting constantly, it's a loose cable or something, I asked DC to work on it but I think it got lost in all the communication. I'll bump it up since I didn't get a chance to visit myself. We might have to upgrade this to the front of the queue as our monitoring system actually also stopped hearing back from it last night. I'm hesitant to provide any updates on network status page at this point even though it's cleaned up a little in terms of our ability because it ends up burying the other mass reports, last time we forgot to bump those back up and it caused a heavy ticket load since it wasn't up top but I'll try to at least resume sending out emails.
@VirMach said: SJCZ004 unfortunately still down, I was going to fly in yesterday and just literally press the button then fly back but flight times didn't work out with the required 24 hour notice to access the site.
If DC remote hands can't press a button then maybe that location isn't working out.
@VirMach said: SJCZ004 unfortunately still down, I was going to fly in yesterday and just literally press the button then fly back but flight times didn't work out with the required 24 hour notice to access the site.
If DC remote hands can't press a button then maybe that location isn't working out.
Yeah that's already been established but at this point I've basically made myself step back and cool off, otherwise I feel like we're going to be migrating everyone around for the next 2 years before we can settle down with an appropriate set of partners that meet or exceed our bare minimum expectations. I had a whole story written out the other day but it got very ranty so I deleted it and didn't post it. The gist of it was that it seems like the entire industry is just pretty screwed up and understaffed/overworked right now or we've got insanely bad luck and there's no way to tell just based on a company's previous reputation that things like this can't happen. Even within the same company, there's a huge variance location to location.
Luckily the locations we still have left with a lot of open space to fill are with xTom, which has so far, even including the more recent cabinets we got with another company, been the clear winner. So we need to just get through this one last hurdle and once I can get down the logistics, I don't mind living in hotels for the next few years.
But yes, it's still very scary, the thought that we could actually be left in a situation (and already have been) where a server can just go down for 1-2 weeks over something so simple.
I noticed that my BF special 2020 doesn't have a Ryzen Migrate button and in it's billing panel, only CentOS 6.8/7, Debian 8.2/9.1 ISOs are available. It has other issues as well, but no problem because I'm idling it. https://lowendspirit.com/discussion/comment/93153/#Comment_93153
Holy crap, we're actually so doomed in Dallas @FrankZ they finally got back to us and said (Flexential) due to the "complexity" of the request, the request being:
Clear CMOS
Reset BMC
I think maybe I made it sound so complicated by asking them to also verify at the end. I'll figure something out for this location, JFC.
Hivelocity's the one that acquired Incero right? I don't know why they didn't come to mind (HV) when I was thinking about Dallas. We can't do it right now but there's absolutely no way we can stay with Flexential at this state, WTF.
@VirMach said:
Hivelocity's the one that acquired Incero right? I don't know why they didn't come to mind (HV) when I was thinking about Dallas. We can't do it right now but there's absolutely no way we can stay with Flexential at this state, WTF.
Yeah, know anyone in Dallas with a rack in their house? They'd be more useful & have better uptime probably.
That's just scary bad. With the headaches you've had and how tired of it all that you must be, I am impressed you didn't scream so loud at them that their eardrums bled
@VirMach said:
Hivelocity's the one that acquired Incero right? I don't know why they didn't come to mind (HV) when I was thinking about Dallas. We can't do it right now but there's absolutely no way we can stay with Flexential at this state, WTF.
I have already moved two of the three VMs I had with you in Dallas to other locations, so I am now safe from future Flexential F'ups.
I hope your contract is not too long with them given what they have shown as their level of competency.
I have no comment on Hivelocity other than to say they have a real nice network in Dallas.
Lol Flexential has one location at the Infomart (1950 N. Stemmons Freeway). I used to work in Dallas many years ago and we had servers at the Infomart in Broadwing (at the time, Level 3 I think bought them?) on 5th or 6th floor (I forget, it was long time ago lol) and a small private space on one of the other floors where I spent most of my time.
I only have one VM in dallas at this point but was debating moving another there.. Think i'll wait on that lol
@VirMach
First, the good news:
My Chicago VPS migrated to Ryzen Chicago. With a simple network reconfiguration in SolusVM, everything went smoothly. Yippee! Now got my tertiary nameserver back up and running. Thanks muchly for the upgrade.
Now, the bad news:
Once again ATL is having issues, after being fine for a week or two. This morning, node 7 if I remember correctly, went down and is dead to the Client Area. "Ach no probs.", I thought, just move my (secondary) nameserver back to the already setup VPS on ATLZ005. Started to go well but began to have connectivity issues with DNS. Hmm. It appears that another VM has nicked my IP address. :'( (23.147.xxx.0 subnet)
Trying the Ryzen Fix IP solution.. (goes off to do other stuff for 1/4hour.)
And the saga continues, eh? ..
[Edited for typos.]
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural. It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
@AlwaysSkint said: Now, the bad news:
Once again ATL is having issues, after being fine for a week or two. This morning, node 7 if I remember correctly, went down and is dead to the Client Area. "Ach no probs.", I thought just move my (secondary) nameserver back to the already setup VPS on ATLZ005. Started to go well but began to have connectivity issues with DNS. Hmm. I appears that another VM has nicked my IP address. (23.147.227.0 subnet)
And the saga continues, eh? ..
Yeah, I saw that as well as a few other things happen and at this point I'm trying to figure out if I actually need to just start a VirBot cloning program so we can ship me off to live in a cage at every datacenter.
We've added an addon for all the people over the last few months that have indicated their extreme displeasure with their service going offline, claiming that as a result they lost thousands. There have also been people that were highly sensitive to IP address changes, location changes, and were in a situation where it was an extremely important production server that required custom solutions and more precise notices and scheduled maintenance windows.
For $800 per month per service we would be able to set up additional monitoring, an account manager with direct contact, and be able to discuss a plan to ensure we can actually minimize downtime on your instance. We'd also be able to effectively in the future ensure that we can meet your customized requirements and make different business decisions for your node. For example, if an IP provider increases costs astronomically, we could still keep it if it's important to you, or we can maintain a node we would have otherwise decommissioned for you.
I don't know if anyone will make this purchase due to the cost, but we basically based the pricing and extrapolated it to where if enough people make a purchase, we could theoretically have a dedicated 24x7 operation of network engineers, system administrators, phone support, covering emergency travel costs, and so on. Basically creating an environment where we could maximize doing absolutely everything physically possible to ensure your production server remains accessible. I've already been thinking about this for some time so I do have an idea of how many people I could theoretically personally support to that level, and then I scaled it accordingly to reach the pricing and then made it look nicer by rounding down.
We finally have an option that could genuinely prevent you from being in a situation where you're frustrated as you're losing millions while your service is inaccessible. We'd even keep an almost-live copy of your service ready to deploy immediately.
If you do not purchase this addon because you're in a situation where it's worth less than $25 per day to have your data, communication with us, or your service online, I do recommend you consider at least doing what you can on your end, such as purchasing a duplicate service, otherwise the default assumption for your let's say $20 a year service will be that it has a value of $20 per year to you and nothing more. This will continue to ensure that our SLA is enough for you and hopefully prevent any situation where requesting SLA credits does not completely solve your issue.
Comments
Okay after working on this all day, outside of the two nodes somehow still down (SJCZ004 and NYCB036) I've counted these specifically and we've gone down from something like 2% of people after the most recent round of migrations to exactly 0.91% of people on Ryzen having a VPS with an incorrect IP or non-bootable. At this point for the other two I basically just need the DC hands to reboot them at least and then I'll probably be able to try to take it from there... I hope it doesn't end up taking two weeks for that type of request but we'll see as we're on our way there.
All the others left have been organized and require a re-migration at this point. This number also does include all the broken VMs from Ryzen migrate button, which make up the majority and are clustered on 28 different nodes. What we're going to do is first create the LVMs manually to fix the majority and boot them up, and then either within 24 hours try to restore the data if it exists, or if it takes any longer, send out a ticket and ask people if they still wanted their old data (in case they begin using it and already loading in their own or just don't have important data.)
This last bit took astronomically longer to identify and fix but now that it's the ending portion of it and heavily organized with all the other problems (the other 1% fixed today) out of the way, it should go smoothly.
Then we'll probably run our auto credit and close ticket script again for everyone, probably piss off a few hundred people at least, but actually be able to get on track with tickets as otherwise the majority of the tickets are going to end up being for issues we already resolved. I'll try to wait on this until we get the other two nodes back up and finalized as well and any remaining migrations done.
From the fact that I still have a number of vps that haven't migrated yet, I'm guessing there is a certain amount total that haven't changed yet. I'm assuming the future migrations left hopefully will be to existing systems that have been tested?
All systems are existing systems that have been tested, some more than others. The newer nodes I've actually spent less time testing, not as a result of being careless, but in that a lot of the issues were already ironed out so it was a quicker fix and it's possible less of these core issues get passed onto those nodes. We have not sold anything on these Ryzen nodes for several weeks now so a good portion of migrations will be to existing nodes that freed up a little bit. The rest will mostly be Hivelocity setups and the final bit will be the last servers I've built which essentially have the best solid state drives (IMO) so they shouldn't "drop off" and the motherboards have been more extensively tested and pre-flashed (mostly.) Any issues these nodes may have will be related to power and brackets falling apart during shipping which get fixed before it goes online. I'm being specific about where I deploy them to ensure if we run into problems, it should hopefully be a quick fix. Only caveat is that the Hivelocity locations will have zero RAM available. We had to do so many RAM swaps and they're all with existing partners, I ran out, but I memory tested everything.
Networking configuration will be the one that gives us no problems straight out the gate for the new nodes as well and I've not sent any additional servers outside of storage to the partner who shall not be named who has taken over a week for a power button press and two weeks for switch configuration or initial node deployments.
We've also gotten quite good at migrations by now so those will hopefully be as smooth as they get for the last bit.
I did also finally get some switch access, so you will either see more nodes up in NYC that we can use for migrations, or you'll see me break the networking.
Is PHXZ004 working fine? I couldn't log in via SSH, so I logged in via VNC, but even ping to xxx.xxx.xxx.1 failed.
No reply required. Instead of writing a reply to me, you can rest your eyes and body 😉
I was going to get a flight to San Jose for tomorrow morning but I'm still waiting to see how we can even get DC access since it's the first time. I might still go on Tuesday and see if I can beat the DC hands to it for SJCZ004 at this point, only about an hour flight.
Phoenix actually has some of our infrastructure on it and I can confirm it's being terrible. But it's still mostly usable. We need to redo network configuration on it still.
all my VPS (LA) can't ping to 8.8.8.8. any explanation on that?
Very unlucky user or very lucky user based on preference, when it happened, etc. Maybe your services are getting migrated to Ryzen, maybe they're all broken.
NYC storage node, I had to do a lot of weird setups for this one to make it work as we faced problems such as a disk going missing, the wrong switch, and so on in the background. Right now it's LACP aggregated 3Gbps.
Originally was supposed to have 36 disks, then had 31, then Amazon delayed it, then either one died or DC lost it so we had 30 and I can't keep waiting at this point to try to do 36 disk RAID. This is a beefier card, I figured we give it a try with two 16 disk arrays to essentially emulate two Tokyo smaller servers or to add further fun challenges if I'm wrong about this decision. I configured dm-cache but I don't know what to do with it yet outside of just maybe using it for a test VM. SolusVM surprisingly added this feature in V2, whenever that ends up happening we'll be ready and you can get a whopping 7GB of Gen4 NVMe cache with your 1TB storage (I can probably get it it up higher it has open slots.) Don't know if it's any good.
Anyway, Geekbench:
https://browser.geekbench.com/v5/cpu/16224250/claim?key=836554
This is a Threadripper 3960X. It's the infamous motherboard I dropped, I killed two of the RAM slots so that's why it's at 192GB instead of the planned 256GB and why 36 disk --> 2 x 16 (32) makes more sense, but I ran a lot of tests on it and it was otherwise healthy.
This 9560 has a weird battery (CacheVault) non-issue, I think it just takes longer to boot or has a more strict requirement before initiating. It's been like that since brand new and I've looked into it in every way I could with the limited time, but it only happens on initial boot and shows healthy state.
Fun fact though, it's now impossible to get a replacement motherboard for these, fully. I've looked so hard. So if it ever dies we'll just do a board swap with Epyc, 5950X, 6950X by then, or something like that.
9560-16i RAID controller.
Oh and we already have like 3 NICs now to make it 10Gbps, 20Gbps and later 40Gbps even maybe. We basically spam ordered and crossed fingers for this one since it's initially a lot larger than Tokyo and needs the 10Gbps. They just need to connect it.
I'm about to send the maintenance email for one of the current storage nodes, while we were running backups this week two of the disks on it died and it's degraded. That's one reason I was rushing this out in the last 2 days, I am absolutely not getting DC hands involved at the old DC, every single RAID-related thing they've done recently has been guaranteed destroying all data and the last time they worked on this they almost jumbled the data. Luckily that means we already have like 80% of a massive storage server already backed up to another storage server and I'll be moving off the people that didn't get backed up first.
YABS on the first VPS:
Basic System Information: --------------------------------- Uptime : 0 days, 0 hours, 4 minutes Processor : QEMU Virtual CPU version 2.5+ CPU cores : 2 @ 3799.952 MHz AES-NI : ❌ Disabled VM-x/AMD-V : ❌ Disabled RAM : 2.9 GiB Swap : 3.0 GiB Disk : 913.2 GiB Distro : Debian GNU/Linux 10 (buster) Kernel : 4.19.0-6-amd64 fio Disk Speed Tests (Mixed R/W 50/50): --------------------------------- Block Size | 4k (IOPS) | 64k (IOPS) ------ | --- ---- | ---- ---- Read | 22.18 MB/s (5.5k) | 7.53 MB/s (117) Write | 22.20 MB/s (5.5k) | 7.94 MB/s (124) Total | 44.38 MB/s (11.0k) | 15.47 MB/s (241) | | Block Size | 512k (IOPS) | 1m (IOPS) ------ | --- ---- | ---- ---- Read | 86.89 MB/s (169) | 116.46 MB/s (113) Write | 91.50 MB/s (178) | 124.21 MB/s (121) Total | 178.39 MB/s (347) | 240.67 MB/s (234) iperf3 Network Speed Tests (IPv4): --------------------------------- Provider | Location (Link) | Send Speed | Recv Speed | | | Clouvider | London, UK (10G) | 213 Mbits/sec | 920 Mbits/sec Online.net | Paris, FR (10G) | 882 Mbits/sec | 744 Mbits/sec Hybula | The Netherlands (40G) | 767 Mbits/sec | 1.75 Gbits/sec Uztelecom | Tashkent, UZ (10G) | 793 Mbits/sec | 499 Mbits/sec Clouvider | NYC, NY, US (10G) | 941 Mbits/sec | 2.82 Gbits/secOkay let's pretend I didn't forget to enable the cache.
fio Disk Speed Tests (Mixed R/W 50/50): --------------------------------- Block Size | 4k (IOPS) | 64k (IOPS) ------ | --- ---- | ---- ---- Read | 313.06 MB/s (78.2k) | 1.18 GB/s (18.4k) Write | 313.88 MB/s (78.4k) | 1.18 GB/s (18.5k) Total | 626.95 MB/s (156.7k) | 2.37 GB/s (37.0k) | | Block Size | 512k (IOPS) | 1m (IOPS) ------ | --- ---- | ---- ---- Read | 5.34 GB/s (10.4k) | 5.11 GB/s (4.9k) Write | 5.63 GB/s (11.0k) | 5.45 GB/s (5.3k) Total | 10.97 GB/s (21.4k) | 10.56 GB/s (10.3k)Alright now let's fill it and bring those numbers down. I'll also try to start moving the other one less aggressively to LAX storage and shifting our disaster recovery backups from that to non-RAID since they're mostly complete. Once we finish the rest of the NY storage we'll let you move back and I apologize in advance for the double move but it's better than it getting yoinked offline.
Any space left on this one for new sign ups?
(Asking for a friend)
blog archives
What is the consensus with networking in Phoenix (PHXZ003), are whole nodes having issues? I cannot ping to 8.8.8.8 from the VPS after trying the usual stuff like a hard power-off then on. I have the "Fix Ryzen IP" option, but don't want to just try stuff randomly if the whole site is having issues.
my LAX had 8.8.8.8 issues for a while , but went back up.
funny enough i was trying to do a bench.sh test and it was coincidentally down, and up afterwards.
I bench YABS 24/7/365 unless it's a leap year.
Not just you, I'm on PHXZ002 and the network was so so before, but now no network at all for last ~24 hours. I don't think you can fix this yourself. VirMach mentioned issues in Phoenix above.
I'm wondering if its LA10GKVM14 since that node has been down almost one month already
I expect he was referring to the shared hosting server in Buffalo with the comment you quoted. I could be wrong, but that was my understanding.
Honestly at this point I think 8.8.8.8 issue is actually not on any specific datacenter's end. We had it happen on our WHMCS for a little bit, and that one's not even hosted with any datacenter we use right now for these Ryzens. We've also had it reported on LAX which is QN and Phoenix which is PhoenixNAP IIRC.
So it has to be something upstream probably unless it's in some weird way related?
LA10GKVM14 was part of a massive PDU or power circuit event with them, where like a dozen plus servers went offline and online, and then offline, causing a bunch of power supplies to fail. A lot of times when we've had power issues like that, CC also moves them to another switch without telling us and then does not configure the networking properly. This is one of those cases, so it's basically been left without functional networking and no one's willing to help. We're marking it at a loss at this point because it's impossible to get them to do anything anymore. It's possible this also has other issues like a failed controller/data corruption as a result of it taking so long for them to hook up another PSU, I think it took 4 or 5 days and the battery drained for the cache but it doesn't matter. We have to try to locate backups which we haven't had luck with so far because backups also failed and were closely tied to the power event, I think we have partial backups.
So right now we're mostly stuck on regenerating services.
Phoenix is completely trashed right now.
i think so too. something related geographically in in US
I bench YABS 24/7/365 unless it's a leap year.
Yeah I don't know how reliable this random site is that I found but:
I think most of us/customers would be happy with fresh idlers servers. Specially after a month of downtime so I wouldn't spend any extra time for hunting backups but I usually to keep my own backups.
Could you please migrate me to Tokyo?
It's been offline for ten days on SJZ004 and the control pannel page timeout
order is: 4678913601
How do you propose a migration of an offline server?
Can you migrate my broken VPS in Phoenix to
getRandomLocation()without data?(Yes, needless to say, this is a joke. I'd rather wait than open a ticket)
there is no important data ,a new vps is enough
I'd actually accept these kinds of request if it was possible to keep it clean but realistically that means you get a new VM and then the old one just kind of hangs around while it's offline and takes up space until we manually verify which ones are abandoned and clear up the space.
For Phoenix though it's technically online so we could start allowing Ryzen migrations. Let me just finish up with what I need to do this morning and I'll activate the button again.
I think it would help some users, but "(non-Phoenix location) has no buttons!!!" and that could open a ticket.
I am not really in a hurry and can wait. That said, I respect your decision.
@Virmach
Additional 7 IPs to VPS two months ago. Some additional IPs have been unavailable for about a month. The additional IPs of the billing panel also cannot be displayed. I opened a ticket on July 6th. I'd appreciate it if you could help me check it out.
Ticket #218767
Button updated.
I can't get it to work but let's see if anyone else has any luck (edit -- looks like it did work for at least one person so far.) To be far I was trying to break it by being impatient and refreshing/closing it. Let me allow it to run for 5 minutes and see what happens instead.
Oh wait I think it's because I activated our trap card
Would it be possible to show that button also ppl with servers in broken nodes (like LA10GKVM14)? That way users could regenerate their services in different nodes
There's no way to differentiate between fully broken and partially broken nodes to display it only to those people.
Did I miss the Ryzen migrate button for Phoenix or was that not a thing?
EDIT: Never mind Phoenix network seems to be working now, slow, but working.
I actually migrated out of a broken Chicago node. I have two machines as a result just one is unavailable.
Took it away while we do more migrations, we can't have it drastically change too much or it'll put the plans out of whack. It'll re-enable later tonight.
Finally got back the physical port numbers for the servers I needed, NYCB036 (formerly 101) should be fixed soon, just need to wrap up what I was working on and configure it in the switch. SJCZ004 unfortunately still down, I was going to fly in yesterday and just literally press the button then fly back but flight times didn't work out with the required 24 hour notice to access the site.
DALZ009 appears to be having problems since at least last night, didn't get a chance to send out emails or look into it yet. LAXA014 still keeps rebooting constantly, it's a loose cable or something, I asked DC to work on it but I think it got lost in all the communication. I'll bump it up since I didn't get a chance to visit myself. We might have to upgrade this to the front of the queue as our monitoring system actually also stopped hearing back from it last night. I'm hesitant to provide any updates on network status page at this point even though it's cleaned up a little in terms of our ability because it ends up burying the other mass reports, last time we forgot to bump those back up and it caused a heavy ticket load since it wasn't up top but I'll try to at least resume sending out emails.
Re-enabled.
Chicago going up today.
Any additional details about datacenter&networks?
If DC remote hands can't press a button then maybe that location isn't working out.
Yeah that's already been established but at this point I've basically made myself step back and cool off, otherwise I feel like we're going to be migrating everyone around for the next 2 years before we can settle down with an appropriate set of partners that meet or exceed our bare minimum expectations. I had a whole story written out the other day but it got very ranty so I deleted it and didn't post it. The gist of it was that it seems like the entire industry is just pretty screwed up and understaffed/overworked right now or we've got insanely bad luck and there's no way to tell just based on a company's previous reputation that things like this can't happen. Even within the same company, there's a huge variance location to location.
Luckily the locations we still have left with a lot of open space to fill are with xTom, which has so far, even including the more recent cabinets we got with another company, been the clear winner. So we need to just get through this one last hurdle and once I can get down the logistics, I don't mind living in hotels for the next few years.
But yes, it's still very scary, the thought that we could actually be left in a situation (and already have been) where a server can just go down for 1-2 weeks over something so simple.
I noticed that my BF special 2020 doesn't have a Ryzen Migrate button and in it's billing panel, only CentOS 6.8/7, Debian 8.2/9.1 ISOs are available. It has other issues as well, but no problem because I'm idling it.
https://lowendspirit.com/discussion/comment/93153/#Comment_93153
Holy crap, we're actually so doomed in Dallas @FrankZ they finally got back to us and said (Flexential) due to the "complexity" of the request, the request being:
I think maybe I made it sound so complicated by asking them to also verify at the end. I'll figure something out for this location, JFC.
Hivelocity's the one that acquired Incero right? I don't know why they didn't come to mind (HV) when I was thinking about Dallas. We can't do it right now but there's absolutely no way we can stay with Flexential at this state, WTF.
Yeah HiVelocity bought Incero. https://www.hivelocity.net/blog/hivelocity-acquires-dallas-iaas-provider-incerocom/
Yeah, know anyone in Dallas with a rack in their house? They'd be more useful & have better uptime probably.
That's just scary bad. With the headaches you've had and how tired of it all that you must be, I am impressed you didn't scream so loud at them that their eardrums bled
https://www.tomshardware.com/news/japanese-government-invests-dollar680-million-in-kioxia-wd-fab
New NVMe for new builds?
I bench YABS 24/7/365 unless it's a leap year.
I have already moved two of the three VMs I had with you in Dallas to other locations, so I am now safe from future Flexential F'ups.
I hope your contract is not too long with them given what they have shown as their level of competency.
I have no comment on Hivelocity other than to say they have a real nice network in Dallas.
Lol Flexential has one location at the Infomart (1950 N. Stemmons Freeway). I used to work in Dallas many years ago and we had servers at the Infomart in Broadwing (at the time, Level 3 I think bought them?) on 5th or 6th floor (I forget, it was long time ago lol) and a small private space on one of the other floors where I spent most of my time.
I only have one VM in dallas at this point but was debating moving another there.. Think i'll wait on that lol
@VirMach
First, the good news:
My Chicago VPS migrated to Ryzen Chicago. With a simple network reconfiguration in SolusVM, everything went smoothly. Yippee! Now got my tertiary nameserver back up and running. Thanks muchly for the upgrade.
Now, the bad news:
Once again ATL is having issues, after being fine for a week or two. This morning, node 7 if I remember correctly, went down and is dead to the Client Area. "Ach no probs.", I thought, just move my (secondary) nameserver back to the already setup VPS on ATLZ005. Started to go well but began to have connectivity issues with DNS. Hmm. It appears that another VM has nicked my IP address. :'( (23.147.xxx.0 subnet)
Trying the Ryzen Fix IP solution.. (goes off to do other stuff for 1/4hour.)
And the saga continues, eh? ..
[Edited for typos.]
In stasis until the shitposting stops/abates.
Than=compare;then=sequence:brought=bring;bought=buy:staffs=pile of sticks:informations/infos=no plural.
It wisnae me! A big boy done it and ran away. || NVMe2G for life! until death (the end is nigh).
Yeah, I saw that as well as a few other things happen and at this point I'm trying to figure out if I actually need to just start a VirBot cloning program so we can ship me off to live in a cage at every datacenter.
Announcement:
We've added an addon for all the people over the last few months that have indicated their extreme displeasure with their service going offline, claiming that as a result they lost thousands. There have also been people that were highly sensitive to IP address changes, location changes, and were in a situation where it was an extremely important production server that required custom solutions and more precise notices and scheduled maintenance windows.
For $800 per month per service we would be able to set up additional monitoring, an account manager with direct contact, and be able to discuss a plan to ensure we can actually minimize downtime on your instance. We'd also be able to effectively in the future ensure that we can meet your customized requirements and make different business decisions for your node. For example, if an IP provider increases costs astronomically, we could still keep it if it's important to you, or we can maintain a node we would have otherwise decommissioned for you.
I don't know if anyone will make this purchase due to the cost, but we basically based the pricing and extrapolated it to where if enough people make a purchase, we could theoretically have a dedicated 24x7 operation of network engineers, system administrators, phone support, covering emergency travel costs, and so on. Basically creating an environment where we could maximize doing absolutely everything physically possible to ensure your production server remains accessible. I've already been thinking about this for some time so I do have an idea of how many people I could theoretically personally support to that level, and then I scaled it accordingly to reach the pricing and then made it look nicer by rounding down.
We finally have an option that could genuinely prevent you from being in a situation where you're frustrated as you're losing millions while your service is inaccessible. We'd even keep an almost-live copy of your service ready to deploy immediately.
If you do not purchase this addon because you're in a situation where it's worth less than $25 per day to have your data, communication with us, or your service online, I do recommend you consider at least doing what you can on your end, such as purchasing a duplicate service, otherwise the default assumption for your let's say $20 a year service will be that it has a value of $20 per year to you and nothing more. This will continue to ensure that our SLA is enough for you and hopefully prevent any situation where requesting SLA credits does not completely solve your issue.