GNU social JP
  • FAQ
  • Login
GNU social JPは日本のGNU socialサーバーです。
Usage/ToS/admin/test/Pleroma FE
  • Public

    • Public
    • Network
    • Groups
    • Featured
    • Popular
    • People

Conversation

Notices

  1. Embed this notice
    Kelly Shortridge (shortridge@hachyderm.io)'s status on Saturday, 03-Feb-2024 09:09:42 JST Kelly Shortridge Kelly Shortridge

    in case there are other nerds out there who haven’t yet read this classic, behold “the case of the 500-mile email” https://www.ibiblio.org/harris/500milemail.html

    I adore the “absurd computer-borne mysteries” genre and kindly ask for more content from the annals of y’all’s careers

    In conversation Saturday, 03-Feb-2024 09:09:42 JST from hachyderm.io permalink

    Attachments


    • clacke likes this.
    • Matthew Lyon, Jesse 🇫🇷 and clacke repeated this.
    • Embed this notice
      Hazel Weakly (hazelweakly@hachyderm.io)'s status on Sunday, 04-Feb-2024 05:53:13 JST Hazel Weakly Hazel Weakly
      in reply to

      @shortridge Actually, I take it back, this is my favorite bug story of all time

      https://www.teamten.com/lawrence/writings/coding-machines/

      The writing is incredible, the twist is beyond absurd. The full implications of it are profound and potentially disturbing

      It's not a read for the lighthearted, it'll take a while, but it's absolutely worth it

      Imagine "reflections on trusting trust" but rendered real and haunting

      In conversation Sunday, 04-Feb-2024 05:53:13 JST permalink
    • Embed this notice
      Wolf480pl (wolf480pl@mstdn.io)'s status on Tuesday, 20-Feb-2024 14:14:01 JST Wolf480pl Wolf480pl
      in reply to
      • Hazel Weakly
      • The Uberduck

      @uberduck @hazelweakly @shortridge
      For me, the OG "reflections on trusting trust" is much more concerning than this. But I guess that's hardly a consolation.

      In conversation Tuesday, 20-Feb-2024 14:14:01 JST permalink
      Haelwenn /элвэн/ :triskell: likes this.
    • Embed this notice
      Wolf480pl (wolf480pl@mstdn.io)'s status on Tuesday, 20-Feb-2024 14:14:02 JST Wolf480pl Wolf480pl
      in reply to
      • Hazel Weakly
      • The Uberduck

      @uberduck @hazelweakly @shortridge
      It's fiction. The timing with the mailman is too convenient from a narrative perspective.

      In conversation Tuesday, 20-Feb-2024 14:14:02 JST permalink
      Haelwenn /элвэн/ :triskell: likes this.
    • Embed this notice
      The Uberduck (uberduck@hachyderm.io)'s status on Tuesday, 20-Feb-2024 14:14:02 JST The Uberduck The Uberduck
      in reply to
      • Wolf480pl
      • Hazel Weakly

      @wolf480pl @hazelweakly @shortridge Okay, but you do get how hanging the fact/fiction decision on that instead of any of the technical details doesn't make me feel any better, right?

      In conversation Tuesday, 20-Feb-2024 14:14:02 JST permalink
    • Embed this notice
      The Uberduck (uberduck@hachyderm.io)'s status on Tuesday, 20-Feb-2024 14:14:03 JST The Uberduck The Uberduck
      in reply to
      • Hazel Weakly

      @hazelweakly @shortridge I got way too far in this before the possibility that it was fiction occurred to me.

      It's fiction. Right?

      Right?

      In conversation Tuesday, 20-Feb-2024 14:14:03 JST permalink

      Attachments


    • Embed this notice
      Haelwenn /элвэн/ :triskell: (lanodan@queer.hacktivis.me)'s status on Tuesday, 20-Feb-2024 14:17:22 JST Haelwenn /элвэн/ :triskell: Haelwenn /элвэн/ :triskell:
      in reply to
      @shortridge One story in this style I really like is https://patrickthomson.tumblr.com/post/2499755681/the-best-debugging-story-ive-ever-heard which I'd dub floor tiles vs. mainframe.
      In conversation Tuesday, 20-Feb-2024 14:17:22 JST permalink

      Attachments

      1. Domain not in remote thumbnail source whitelist: 64.media.tumblr.com
        The Best Debugging Story I've Ever Heard
        from djinn and juice
        Back in the early 80's, my dad worked at Storage Technology, a now-defunct corporate entity that made tape drives and pneumatic systems to drive these tapes at high speeds – for that period of...
    • Embed this notice
      Haelwenn /элвэн/ :triskell: (lanodan@queer.hacktivis.me)'s status on Tuesday, 20-Feb-2024 14:20:42 JST Haelwenn /элвэн/ :triskell: Haelwenn /элвэн/ :triskell:
      in reply to
      • Haelwenn /элвэн/ :triskell:
      @shortridge And I think you'd also enjoy https://www.ecb.torontomu.ca/~elf/hack/recovery.html even though the problem is known right from the start, how they pieced it together for a recovery is just glorious.
      In conversation Tuesday, 20-Feb-2024 14:20:42 JST permalink

      Attachments

      1. No result found on File_thumbnail lookup.
        https://www.ecb.torontomu.ca/~elf/hack/recovery.html
    • Embed this notice
      Resuna (resuna@ohai.social)'s status on Thursday, 28-Mar-2024 01:31:41 JST Resuna Resuna
      in reply to

      @shortridge I remember that one, by the last paragraph where he used the word "millilightseconds", at that point the vague "this seems familiar" feeling collapsed into "yes I've read this before". Something something quantum something.

      In conversation Thursday, 28-Mar-2024 01:31:41 JST permalink
      clacke likes this.
    • Embed this notice
      maswan (maswan@mastodon.acc.sunet.se)'s status on Thursday, 28-Mar-2024 01:31:46 JST maswan maswan
      in reply to

      @shortridge
      I had a server (back when servers came in towers, not rack units) that would lock up hard randomly at about weekly frequency, unless there was a PS/2 mouse plugged in. We put it ziptied in a couple of unused 5.25" bays.

      This server spent years after that as the distributor of Debian to European mirrors.

      The cause? Some memory mapping bug in bios, we think.

      In conversation Thursday, 28-Mar-2024 01:31:46 JST permalink
      clacke likes this.
    • Embed this notice
      DCoder 🇱🇹❤🇺🇦 (dcoderlt@ohai.social)'s status on Thursday, 28-Mar-2024 01:31:51 JST DCoder 🇱🇹❤🇺🇦 DCoder 🇱🇹❤🇺🇦
      in reply to

      @shortridge
      One of the first things I had to do as a professional developer decades ago was figure out why a client’s homemade CMS was periodically losing all content.
      Turns out the admin panel’s access check redirected guests away, but *did not terminate* itself. So a web crawler came in, got the redirect followed by the admin panel dashboard with a list of latest articles, and links to delete.php?article_id=123 next to each one… It dutifully crawled all those, and poof, no more articles.

      In conversation Thursday, 28-Mar-2024 01:31:51 JST permalink
      clacke likes this.
    • Embed this notice
      Zimmie (bob_zim@infosec.exchange)'s status on Friday, 13-Jun-2025 08:07:26 JST Zimmie Zimmie
      in reply to

      @shortridge While working tech support, I got a call on a Monday. Some VPNs which had been working on Friday were no longer working. After a little digging, we found the negotiation was failing due to a certificate validation failure.

      The certificate validation was failing because the system couldn’t check the certificate revocation list (CRL).

      The system couldn’t check the CRL because it was too big. The software doing the validation only allocated 512kB to store the CRL, and it was bigger than that. This is from a private certificate authority, though, and 512kB is a *LOT* of revoked certificates. Shouldn’t be possible for this environment to hit within a human lifespan.

      Turns out the CRL was nearly a megabyte! What gives? We check the certificate authority, and it’s revoking and reissuing every single certificate it has signed once per second.

      The revocations say all the certificates (including the certificate authority’s) are expired. We check the expiration date of the certificate authority, and it’s set to some time in 1910. What? It was around here I started to suspect what had happened.

      The certificate authority isn’t valid before some time in 2037. It was waking up every second, seeing the current date was after the expiration date and reissuing everything. But time is linear, so it doesn’t make sense to reissue an expired certificate with an earlier not-valid-before date, so it reissued all the certs with the same dates and went to sleep. One second later, it woke up and did the whole process over again. But why the clearly invalid dates on the CA?

      The CA operation log was packed with revocations and reissues, but I eventually found the reissues which changed the validity dates of the CA’s certificate. Sure enough, it reissued itself in 2037 and the expiration date was set to 2037 plus ten years, which fell victim to the 2038 limitation. But it’s not 2037, so why did the system think it was?

      The OS running the CA was set to sync with NTP every 120 seconds, and it used a really bad NTP client which blindly set the time to whatever the NTP server gave it. No sanity checking, no drifting. Just get the time, set the time. OS logs showed most of the time, the clock adjustment was a fraction of a second. Then some time on Saturday, there was an adjustment of tens of thousands of seconds forward. The next adjustment was hundreds of thousands of seconds forward. Tens of millions of seconds forward. Eventually it hit billions of seconds backwards, taking the system clock back to 1904 or so. The NTP server was racing forward through the 32-bit timestamp space.

      At some point, the NTP server handed out a date in 2037 which was after the CA’s expiration. It reissued itself as I described above, and a date math bug resulted in a cert which expired before it was valid. So now we have an explanation for the CRL being so huge. On to the NTP server!

      Turns out they had an NTP “appliance” with a radio clock (i.e, a CDMA radio, GPS receiver, etc.). Whoever built it had done so in a really questionable way. It seems it had a faulty internal clock which was very fast. If it lost upstream time for a while, then reacquired it after the internal clock had accumulated a whole extra second, the server didn’t let itself step backwards or extend the duration of a second. The math it used to correct its internal clock somehow resulted in dramatically shortening the duration of a second until it wrapped in 2038 and eventually ended up at the correct time.

      Ultimately found three issues:
      • An OS with an overly-simplistic NTP client
      • A certificate authority with a bad date math system
      • An NTP server with design issues and bad hardware

      Edit: The popularity of this story has me thinking about it some more.

      The 2038 problem happens because when the first bit of a 32-bit value is 1 and you use it as a signed integer, it’s interpreted as a negative number in 2’s complement representation. But C has no protection from treating the same value as signed in some contexts and unsigned in others. If you start with a signed 32-bit integer with the value -1, it is represented in memory as 0xFFFFFFFF. If you then use it as an unsigned integer, it becomes the value 4,294,967,296.

      I bet the NTP box subtracted the internal clock’s seconds from the radio clock’s seconds as signed integers (getting -1 seconds), then treated it as an unsigned integer when figuring out how to adjust the tick rate. It suddenly thought the clock was four billion seconds behind, so it really has to sprint forward to catch up!

      In my experience, the most baffling behavior is almost always caused by very small mistakes. This small mistake would explain the behavior.

      In conversation about a year ago permalink

      Attachments


      Haelwenn /элвэн/ :triskell: and clacke like this.
      Haelwenn /элвэн/ :triskell: repeated this.
    • Embed this notice
      Zimmie (bob_zim@infosec.exchange)'s status on Friday, 13-Jun-2025 16:22:59 JST Zimmie Zimmie
      in reply to

      @shortridge Some time later, I was no longer working tech support. I got hired to do network and firewall stuff for a fairly large company. At one point, they decided to relocate the office where a lot of the operations and monitoring staff worked. They moved the whole application monitoring team to the new building with the unproven infrastructure first, because some people in charge made very bad decisions.

      The monitoring team gets to the new building, and they can’t access any of their monitoring systems. Clearly a problem with the new office, right? They go through a few environments to get to their monitoring systems, so I log in to the remote access VPN for the first one and confirm the first firewall they hit sees their traffic and isn’t dropping it.

      I go to log in to the remote access VPN for the second environment, where the monitoring systems actually live. I’m able to start the connection, but it never prompts me for my credentials, and the tunnel never comes up. Huh. That’s weird.

      Well, I’ll just get in through the DR version of the second environment. Connection works and it prompts me for my credentials, but it rejects them. I try again, in case I made a mistake entering the passphrase for my key, but it’s still rejected. Huh. That’s weird.

      I eventually find a working way in. I’m able to ping all the relevant systems, I’m able to make TCP connections via telnet, but trying to actually use a service like SSH or MSRDP just hangs. But wait! I can connect to my firewalls via SSH! So what’s common among the broken systems?

      All the broken systems are VMs. I start testing connections to other things which I know are VMs. They all behave the same. Ping works, TCP connections work, but data over the connections gets no response.

      I bring in the virtualization team. Some of us drive in to the datacenter hosting the VMs giving us trouble. Someone quickly realizes the single SAN hosting all of the VMs’ drives was up, but wasn’t responding to storage requests. Effectively the drive had been pulled out of every single VM. Now we have an explanation for why all the VMs seem to be broken.

      With most operating systems, the network stack is wired in RAM and can’t be swapped out. The network stack handles responding to pings and opening TCP connections on listening ports. Once a TCP connection is opened, it requests a copy of the listening service from storage to handle the connection. With storage no longer responding, the network stack never gets the copy of the service to handle the connection, so data doesn’t work.

      Why couldn’t I connect to the second VPN endpoint? Well, some people in charge made very bad decisions. They had decided that since VMs are the future, the VPN endpoints in that facility should be moved from dedicated hardware to VMs stored on the SAN. They hadn’t gotten to the first VPN endpoint yet, but that environment wasn’t allowed to connect in to the second environment.

      Okay, but I could connect to the other site’s VPN endpoint, and the other site didn’t have any problems. Why didn’t it accept my credentials? Well, some people in charge made very bad decisions (you may be noticing a theme!). All authentication was run through some VMs which were stored on the SAN. The VPN boxes in the working location were set to monitor the health of the authentication boxes in the failed location by pinging them. As long as they responded to ping, they were good, so the VPN boxes wouldn’t fail over to using their local authentication boxes. And a computer with its drive pulled can still respond to ping with just the network stack in RAM.

      Once we realized what was going on, we physically connected to the WAN routers and added routes to prevent the two sites from reaching each other’s authentication boxes. Presto! We could now log in via the DR environment as normal. The other infrastructure teams were then able to start digging into their parts.

      But why is the SAN unresponsive? Turns out this particular SAN vendor had an option for what to do under certain failure conditions: it could fail read-only or fail completely silent. This one was set to fail silent, and it had filled up.

      I wasn’t directly involved in fixing the SAN. I know the manager over the SAN team had been sounding the alarm for months before it filled. I also know there were multiple levels of bad configuration, such as more space offered by LUNs than the SAN could physically provide.

      Big takeaways:
      1. Make sure your access to fix a system doesn’t depend on that system. It’s really easy to accidentally introduce dependency cycles, and it takes constant work to avoid them.
      2. Superficial tests like whether you can ping something can’t detect some pretty major failures. More significant tests are more likely to notice the problem.
      3. When something is critical to an environment, maybe have more than one of them? The SAN had internal redundancy to deal with faulty drives and so on, but all the storage was in one giant pool. Multiple SAN systems can provide a bulkhead such that breaking one would not break all VMs.

      In conversation about a year ago permalink
      Haelwenn /элвэн/ :triskell: and clacke like this.
    • Embed this notice
      Stephen Paulger (aimaz@mstdn.social)'s status on Saturday, 14-Jun-2025 20:15:42 JST Stephen Paulger Stephen Paulger
      in reply to

      @shortridge https://500mile.email a few of the classics are on this site named after the all time best one.

      In conversation about a year ago permalink

      Attachments

      1. No result found on File_thumbnail lookup.
        500 Mile Email
        500 Mile Email - Absurd Software Bug Stories
      clacke likes this.
    • Embed this notice
      Janis La Couvée (lacouvee@mastodon.online)'s status on Saturday, 14-Jun-2025 20:15:44 JST Janis La Couvée Janis La Couvée
      in reply to
      • Kensan

      @Kensan @shortridge is this a serious bug? Because when my Office 2013 bites the dust I'll be moving to Open Office/Libre Office and I have a Brother printer - just sayin'.

      In conversation about a year ago permalink
    • Embed this notice
      Kensan (kensan@mastodon.social)'s status on Saturday, 14-Jun-2025 20:15:44 JST Kensan Kensan
      in reply to
      • Janis La Couvée

      @lacouvee The bug has been fixed in 2009. ;)

      @shortridge

      In conversation about a year ago permalink

      Attachments


      1. https://files.mastodon.social/media_attachments/files/111/869/793/264/398/165/original/35879de63acc4006.png
      clacke likes this.
    • Embed this notice
      Kensan (kensan@mastodon.social)'s status on Saturday, 14-Jun-2025 20:15:45 JST Kensan Kensan
      in reply to

      @shortridge Have you heard of the “OpenOffice.org won’t print to Brother printers on Tuesdays (but works on other days of the week)” bug?

      http://catless.ncl.ac.uk/Risks/25/77#subj14.1

      https://mdzlog.alcor.net/2009/08/15/bohrbugs-openoffice-org-wont-print-on-tuesdays/

      Ubuntu bug:
      https://bugs.launchpad.net/ubuntu/+source/file/+bug/248619

      In conversation about a year ago permalink

      Attachments

      1. Domain not in remote thumbnail source whitelist: bugs.launchpad.net
        Bug #248619 “file incorrectly labeled as Erlang JAM file (OOo do...” : Bugs : file package : Ubuntu
        Binary package hint: file anon@x-X:~$ echo "1/2 Tue" >> file && file file file: Jan 22 14:32:44 MET 1991\011Erlang JAM file - version 4.2 anon@x-X:~$ file --version file-4.21 magic file from /etc/magic:/usr/share/file/magic If you look at the magic file: 4 string Tue Jan 22 14:32:44 MET 1991 Erlang JAM file - version 4.2 So at the fourth byte of the file, it's looking for Tue Jan 22 14:32:44 MET 1991 Since the spaces aren't escaped, it sees that the file contains "Tue" at the four...
    • Embed this notice
      Kevin Marks (kevinmarks@xoxo.zone)'s status on Saturday, 14-Jun-2025 20:15:57 JST Kevin Marks Kevin Marks
      in reply to
      • Kensan

      @Kensan @shortridge this reminds me of the "python only parses dates correctly after the 12th of the month" problem I had. (The dates in the files I was being sent had been changed to UK dd/mm/YYYY format. Python assumes mm/dd/YYYY unless the mm>12)

      In conversation about a year ago permalink
      clacke likes this.
    • Embed this notice
      Hugo van Kemenade (hugovk@mastodon.social)'s status on Saturday, 14-Jun-2025 20:16:10 JST Hugo van Kemenade Hugo van Kemenade
      in reply to
      • Kevin Marks
      • Kensan

      @KevinMarks @Kensan @shortridge That's quite the gotcha!

      Well, if you're not using ISO dates, and don't tell it what format is being used, this library has to make some sort of guess between mm/dd/YYYY and dd/mm/YYYY.

      And iirc you have to tell the standard library which date format to parse, it won't guess.

      In conversation about a year ago permalink
      clacke repeated this.
    • Embed this notice
      clacke (clacke@libranet.de)'s status on Saturday, 14-Jun-2025 20:16:10 JST clacke clacke
      in reply to
      • Kevin Marks
      • Hugo van Kemenade
      • Kensan

      Back in the day I wrote some LotusScript to convert a bunch of textual dates, and I implemented a heuristic "some of these we can determine mm/dd vs dd/mm because the dd ≥ 13, and we know they should be in order, so we can use those anchor dates to further narrow down the others". 😅


      @hugovk @KevinMarks @Kensan @shortridge

      In conversation about a year ago permalink
    • Embed this notice
      Hugo van Kemenade (hugovk@mastodon.social)'s status on Saturday, 14-Jun-2025 20:16:12 JST Hugo van Kemenade Hugo van Kemenade
      in reply to
      • Kevin Marks
      • Kensan

      @KevinMarks @Kensan @shortridge Which bit of Python assumes mm/dd/YYYY unless mm>12?

      In conversation about a year ago permalink
    • Embed this notice
      Kevin Marks (kevinmarks@xoxo.zone)'s status on Saturday, 14-Jun-2025 20:16:12 JST Kevin Marks Kevin Marks
      in reply to
      • Hugo van Kemenade
      • Kensan

      @hugovk @Kensan @shortridge https://dateutil.readthedocs.io/en/stable/parser.html see the dayfirst and yearfirst settings docs

      In conversation about a year ago permalink

      Attachments

      1. No result found on File_thumbnail lookup.
        parser — dateutil 3.9.0 documentation
    • Embed this notice
      Ariaflame (ariaflame@masto.ai)'s status on Saturday, 14-Jun-2025 20:16:22 JST Ariaflame Ariaflame
      in reply to
      • Kevin Marks
      • Kensan

      @KevinMarks @Kensan @shortridge This is why all dates should be YYYY-mm-dd

      In conversation about a year ago permalink
      clacke likes this.
    • Embed this notice
      Ariaflame (ariaflame@masto.ai)'s status on Saturday, 14-Jun-2025 20:16:25 JST Ariaflame Ariaflame
      in reply to
      • Kevin Marks
      • Kensan

      @KevinMarks @Kensan @shortridge Why I am glad to live somewhere that doesn't do DST

      In conversation about a year ago permalink
      clacke likes this.
    • Embed this notice
      Kevin Marks (kevinmarks@xoxo.zone)'s status on Saturday, 14-Jun-2025 20:16:26 JST Kevin Marks Kevin Marks
      in reply to
      • Ariaflame
      • Kensan

      @ariaflame @Kensan @shortridge the other challenge with being back in the UK is that half the year UTC and local time are the same.
      Oh, and being near 0 longitude..

      In conversation about a year ago permalink
    • Embed this notice
      Kensan (kensan@mastodon.social)'s status on Saturday, 14-Jun-2025 20:16:35 JST Kensan Kensan
      in reply to

      @shortridge Sorry for resurrecting this Thread but this one belongs here:

      "The Wi-Fi only works when it's raining."

      https://predr.ag/blog/wifi-only-works-when-its-raining/

      In conversation about a year ago permalink

      Attachments

      1. Domain not in remote thumbnail source whitelist: predr.ag
        The Wi-Fi only works when it's raining
        from @PredragGruevski
        The strangest hardware problem I've ever had to debug.
      clacke likes this.
    • Embed this notice
      Michael Bacon (michaeltbacon@social.coop)'s status on Saturday, 14-Jun-2025 20:16:48 JST Michael Bacon Michael Bacon
      in reply to

      @shortridge And yes, one of the annoying things about running Solaris back in the day is that every time you installed an OS patch, it would *remove* Sendmail 8 and put Sendmail 5 BACK ON.

      Like, gee, thanks, Sun.

      In conversation about a year ago permalink
      clacke likes this.
    • Embed this notice
      maswan (maswan@mastodon.acc.sunet.se)'s status on Saturday, 14-Jun-2025 20:16:51 JST maswan maswan
      in reply to
      • Michael Bacon

      @MichaelTBacon Oh yes, the era of the mandatory post-patch-cleanup script to avoid RCE via Solaris' horribly old and exploitable sendmail.

      Much have changed, now we can trust Debian and Ubuntu to ship security patches to the degree that automated updates is good on (most packages for most) production servers.

      @shortridge

      In conversation about a year ago permalink
      clacke likes this.

Feeds

  • Activity Streams
  • RSS 2.0
  • Atom
  • Help
  • About
  • FAQ
  • TOS
  • Privacy
  • Source
  • Version
  • Contact

GNU social JP is a social network, courtesy of GNU social JP管理人. It runs on GNU social, version 2.0.2-dev, available under the GNU Affero General Public License.

Creative Commons Attribution 3.0 All GNU social JP content and data are available under the Creative Commons Attribution 3.0 license.