• Re: Schr”dinger's hash

    From Charles Curley@3:633/10 to All on Sunday, June 21, 2026 23:10:01
    On Fri, 22 May 2026 09:53:17 -0600
    Charles Curley <charlescurley@charlescurley.com> wrote:

    I have four four terabyte hard drives. Each has a partition on it. The
    four partitions comprise a RAID 5 array using mdadm. On top of that,
    LUKS encryption, then LVM with ext4 logical volumes.

    I believe I have found a solution to this problem. I installed the
    backports kernel. Since then I have run more than four hours solid of
    tests and not found a single error.

    I did replace one hard drive. While that resulted in a quieter office,
    it did not solve the problem.

    Checking voltages from the power supply and the wall with a digital
    volt meter did not show any out of spec problems.

    --
    Does anybody read signatures any more?

    https://charlescurley.com
    https://charlescurley.com/blog/

    --- PyGate Linux v1.5.17
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From David Christensen@3:633/10 to All on Monday, June 22, 2026 09:40:01
    On Fri, 22 May 2026 09:53:17 -0600 Charles Curley wrote:
    I have four four terabyte hard drives. Each has a partition on it. The
    four partitions comprise a RAID 5 array using mdadm. On top of that,
    LUKS encryption, then LVM with ext4 logical volumes.

    On one LVM partition I have a number of backup files, tarred,
    bzipped, and sha256 and sha512 summed. I have a script which will find checksum files, and execute the appropriate program to test the
    archives. It puts each program into the background, parallising any
    number of checksum tests.


    On 6/21/26 14:05, Charles Curley wrote:
    I believe I have found a solution to this problem. I installed the
    backports kernel. Since then I have run more than four hours solid of
    tests and not found a single error.

    I did replace one hard drive. While that resulted in a quieter office,
    it did not solve the problem.

    Checking voltages from the power supply and the wall with a digital
    volt meter did not show any out of spec problems.


    I am glad that your storage is working correctly now.


    Please run and post the following commands with both the previous kernel
    and the backports kernel:

    $ cat /etc/debian_version

    $ uname -a


    I will assume your script spawns a separate, isolated process for each checksum file.


    If you have ruled out the power supply, memory, and disks, another
    possibility could be a race condition in the kernel and/or I/O stack
    that is triggered when multiple processes access storage in parallel.


    Are the checksum errors repeatable on another computer with a similar
    storage architecture and the previous kernel? If so, do they disappear
    with the backports kernel?


    David

    --- PyGate Linux v1.5.17
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From Charles Curley@3:633/10 to All on Monday, June 22, 2026 19:00:01
    On Mon, 22 Jun 2026 00:32:19 -0700
    David Christensen <dpchrist@holgerdanske.com> wrote:

    I am glad that your storage is working correctly now.

    Thank you.



    Please run and post the following commands with both the previous
    kernel and the backports kernel:

    $ cat /etc/debian_version

    $ uname -a

    Backport

    root@hawk:~# cat /etc/debian_version
    13.5
    root@hawk:~# cat /etc/os-release
    PRETTY_NAME="Debian GNU/Linux 13 (trixie)"
    NAME="Debian GNU/Linux"
    VERSION_ID="13"
    VERSION="13 (trixie)"
    VERSION_CODENAME=trixie
    DEBIAN_VERSION_FULL=13.5
    ID=debian
    HOME_URL="https://www.debian.org/"
    SUPPORT_URL="https://www.debian.org/support" BUG_REPORT_URL="https://bugs.debian.org/"
    root@hawk:~# uname -s
    Linux
    root@hawk:~# uname -a
    Linux hawk 7.0.10+deb13-amd64 #1 SMP PREEMPT_DYNAMIC Debian 7.0.10-1~bpo13+1 (2026-05-28) x86_64 GNU/Linux
    root@hawk:~#

    Prior kernel is the same except the kernel is

    linux-image-6.12.90+deb13.1-amd64 6.12.90-2 amd64

    I think that this problem first showed up late April or early May. If
    that is correct, the following extracts from my apt logs might help.
    These were copied and pasted unwrapped, so unless they were mangled in
    transit, long lines should be preserved.

    Start-Date: 2026-04-13 12:46:25
    Commandline: apt install -t trixie-backports linux-image-amd64
    Install: linux-image-6.19.10+deb13-amd64:amd64 (6.19.10-1~bpo13+1, automatic), linux-modules-6.19.10+deb13-amd64:amd64 (6.19.10-1~bpo13+1, automatic), linux-base-amd64:amd64 (6.19.10-1~bpo13+1, automatic), linux-binary-6.19.10+deb13-amd64:amd64 (6.19.10-1~bpo13+1, automatic), linux-base-6.19.10+deb13-amd64:amd64 (6.19.10-1~bpo13+1, automatic)
    Upgrade: linux-image-amd64:amd64 (6.12.74-2, 6.19.10-1~bpo13+1), linux-libc-dev:amd64 (6.12.74-2, 6.19.10-1~bpo13+1)
    End-Date: 2026-04-13 12:46:52

    Start-Date: 2026-04-16 08:45:37
    Commandline: apt remove linux-image-amd64
    Remove: linux-image-amd64:amd64 (6.19.10-1~bpo13+1)
    End-Date: 2026-04-16 08:45:39

    Start-Date: 2026-04-16 08:45:47
    Commandline: apt install linux-image-amd64
    Install: linux-image-amd64:amd64 (6.12.74-2)
    End-Date: 2026-04-16 08:45:49

    Start-Date: 2026-04-18 11:27:13
    Commandline: apt install -t trixie-backports linux-image-amd64
    Install: linux-image-6.19.11+deb13-amd64:amd64 (6.19.11-1~bpo13+1,
    automatic), linux-modules-6.19.11+deb13-amd64:amd64 (6.19.11-1~bpo13+1, automatic), linux-binary-6.19.11+deb13-amd64:amd64 (6.19.11-1~bpo13+1, automatic) Upgrade: linux-image-amd64:amd64 (6.12.74-2,
    6.19.11-1~bpo13+1) End-Date: 2026-04-18 11:27:44

    Start-Date: 2026-04-21 10:35:49
    Commandline: apt purge linux-image-amd64
    Purge: linux-image-amd64:amd64 (6.19.11-1~bpo13+1)
    End-Date: 2026-04-21 10:35:52

    Start-Date: 2026-04-21 10:36:06
    Commandline: apt install linux-image-amd64
    Install: linux-image-amd64:amd64 (6.12.74-2)
    End-Date: 2026-04-21 10:36:08


    Start-Date: 2026-05-01 04:40:22
    Commandline: /usr/bin/unattended-upgrade
    Install: linux-image-6.12.85+deb13-amd64:amd64 (6.12.85-1, automatic)
    Upgrade: linux-image-amd64:amd64 (6.12.74-2, 6.12.85-1)
    End-Date: 2026-05-01 04:40:45

    Start-Date: 2026-05-09 04:49:58
    Commandline: /usr/bin/unattended-upgrade
    Install: linux-image-6.12.86+deb13-amd64:amd64 (6.12.86-1, automatic)

    Start-Date: 2026-05-16 04:01:47
    Commandline: /usr/bin/unattended-upgrade
    Install: linux-image-6.12.88+deb13-amd64:amd64 (6.12.88-1, automatic)

    Start-Date: 2026-05-24 11:54:07
    Commandline: /usr/bin/unattended-upgrade
    Install: linux-image-6.12.90+deb13-amd64:amd64 (6.12.90-1, automatic)

    Start-Date: 2026-05-29 04:08:32
    Commandline: /usr/bin/unattended-upgrade
    Install: linux-image-6.12.90+deb13.1-amd64:amd64 (6.12.90-2, automatic)

    Start-Date: 2026-06-20 11:19:40
    Commandline: apt install -t trixie-backports linux-image-amd64

    Apparently, twice in April I had tried the backports kernel(s), and
    found them unsatisfactory. So possibly the fix came in between linux-image-6.19.11+deb13-amd64 and the current backports kernel, linux-image-7.0.10+deb13-amd64. Quite possibly a major version number
    might even inadvertently fix something as subtle as this.



    I will assume your script spawns a separate, isolated process for
    each checksum file.

    Correct. It does a find on *.sha256sums, *.sha512sums, and several other suffixes. It then sliced and dices to figure out the appropriate
    program to call, sha256sum and sha512sum, respectively.

    The files are created using paths relative to the current directory, so
    when the script runs, it will pushd to that directory

    The key line is

    nice "${prog}" "${opts}" -c "${file}" &

    Where $prog is the result of the slicing and dicing, opts='--quiet',
    and $file the file to be scanned.

    I recently added the option to limit the number of background tasks to
    the number of processors (nproc --all). That reduces but does not
    eliminate the number of errors.



    If you have ruled out the power supply, memory, and disks, another possibility could be a race condition in the kernel and/or I/O stack
    that is triggered when multiple processes access storage in parallel.

    I wouldn't say those are all ruled out. But the fact that the
    backport kernel is a major version number change, and that it appears
    to have solved the problem is highly suggestive. With that caveat, I
    concur. That is definitely something for the kernel folks to look at.



    Are the checksum errors repeatable on another computer with a similar storage architecture and the previous kernel? If so, do they
    disappear with the backports kernel?

    I have a much more recent and much faster computer, peregrine, with nvme storage and 12 cores. Hawk has eight cores, and spinning rust. Hawk is
    where the problem has shown up. Peregrine does not show the problem.

    For 18 gig of data, hawk: 2m18.318s, peregrine 0m24.233s.




    David



    --
    Does anybody read signatures any more?

    https://charlescurley.com
    https://charlescurley.com/blog/

    --- PyGate Linux v1.5.17
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From David Christensen@3:633/10 to All on Tuesday, June 23, 2026 00:50:02
    On 6/22/26 09:54, Charles Curley wrote:
    On Mon, 22 Jun 2026 00:32:19 -0700 David Christensen wrote:
    Please run and post the following commands with both the previous
    kernel and the backports kernel:

    $ cat /etc/debian_version

    $ uname -a

    Backport

    root@hawk:~# cat /etc/debian_version
    13.5
    root@hawk:~# cat /etc/os-release
    PRETTY_NAME="Debian GNU/Linux 13 (trixie)"
    NAME="Debian GNU/Linux"
    VERSION_ID="13"
    VERSION="13 (trixie)"
    VERSION_CODENAME=trixie
    DEBIAN_VERSION_FULL=13.5
    ID=debian
    HOME_URL="https://www.debian.org/" SUPPORT_URL="https://www.debian.org/support" BUG_REPORT_URL="https://bugs.debian.org/"
    root@hawk:~# uname -s
    Linux
    root@hawk:~# uname -a
    Linux hawk 7.0.10+deb13-amd64 #1 SMP PREEMPT_DYNAMIC Debian 7.0.10-1~bpo13+1 (2026-05-28) x86_64 GNU/Linux
    root@hawk:~#

    Prior kernel is the same except the kernel is

    linux-image-6.12.90+deb13.1-amd64 6.12.90-2 amd64


    Good information.


    I think that this problem first showed up late April or early May. If
    that is correct, the following extracts from my apt logs might help.
    These were copied and pasted unwrapped, so unless they were mangled in transit, long lines should be preserved.

    ...

    Apparently, twice in April I had tried the backports kernel(s), and
    found them unsatisfactory. So possibly the fix came in between linux-image-6.19.11+deb13-amd64 and the current backports kernel, linux-image-7.0.10+deb13-amd64.


    AIUI if someone can come up with a shell script or program whose exit
    value reliably indicates the presence or absence of a bug, Git can do a
    binary search over a range of commits and locate the commit where the
    bug originated.


    Quite possibly a major version number
    might even inadvertently fix something as subtle as this.


    Agreed.


    I will assume your script spawns a separate, isolated process for
    each checksum file.

    Correct. It does a find on *.sha256sums, *.sha512sums, and several other suffixes. It then sliced and dices to figure out the appropriate
    program to call, sha256sum and sha512sum, respectively.

    The files are created using paths relative to the current directory, so
    when the script runs, it will pushd to that directory

    The key line is

    nice "${prog}" "${opts}" -c "${file}" &

    Where $prog is the result of the slicing and dicing, opts='--quiet',
    and $file the file to be scanned.


    I use a similar workflow for image files and also wrote a script to
    generate and verify checksum files.


    I recently added the option to limit the number of background tasks to
    the number of processors (nproc --all). That reduces but does not
    eliminate the number of errors.


    That is another clue that there is a race condition related to parallel I/O.


    If you have ruled out the power supply, memory, and disks, another
    possibility could be a race condition in the kernel and/or I/O stack
    that is triggered when multiple processes access storage in parallel.

    I wouldn't say those are all ruled out. But the fact that the
    backport kernel is a major version number change, and that it appears
    to have solved the problem is highly suggestive. With that caveat, I
    concur. That is definitely something for the kernel folks to look at.


    Agreed.


    Are the checksum errors repeatable on another computer with a similar
    storage architecture and the previous kernel? If so, do they
    disappear with the backports kernel?

    I have a much more recent and much faster computer, peregrine, with nvme storage and 12 cores. Hawk has eight cores, and spinning rust. Hawk is
    where the problem has shown up. Peregrine does not show the problem.

    For 18 gig of data, hawk: 2m18.318s, peregrine 0m24.233s.


    That clue makes me think the race condition is the SCSI stack.


    It looks like there is a newer kernel for Trixie. It is best to file
    bug reports against current packages. Can you test it?

    https://packages.debian.org/stable/kernel/linux-image-6.12.94+deb13-amd64


    David

    --- PyGate Linux v1.5.17
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From The Wanderer@3:633/10 to All on Tuesday, June 23, 2026 04:20:01
    On 2026-06-22 at 18:40, David Christensen wrote:
    On 6/22/26 09:54, Charles Curley wrote:
    I think that this problem first showed up late April or early May.
    If that is correct, the following extracts from my apt logs might
    help. These were copied and pasted unwrapped, so unless they were
    mangled in transit, long lines should be preserved.

    ...

    Apparently, twice in April I had tried the backports kernel(s),
    and found them unsatisfactory. So possibly the fix came in between
    linux-image-6.19.11+deb13-amd64 and the current backports kernel,
    linux-image-7.0.10+deb13-amd64.

    AIUI if someone can come up with a shell script or program whose exit
    value reliably indicates the presence or absence of a bug, Git can
    do a binary search over a range of commits and locate the commit
    where the bug originated.
    ...assuming that there aren't any commits broken for other reasons, or otherwise commits where the codebase can't be built far enough for the
    script or program to be able to do its thing, that will be hit along the
    way.
    That could be folded in under "reliably", of course - but it's
    sufficiently far out of the scope of what people might be expected to
    think of for that term, if not previously familiar with what such
    bisections can involve, that it seems worth calling out explicitly.
    That said, git can actually do this even *without* such a
    script/program, as long as you're willing and able to test each
    candidate commit manually. The use of a script or program to automate it
    is actually a subset of the functionality of the 'git bisect'
    sub-command, specifically the sub-sub-command 'git bisect run'; see 'git
    help bisect' for the documentation, most of which is about the version
    of the process where the validation of each commit is done manually
    rather than by a script.
    --
    The Wanderer
    The reasonable man adapts himself to the world; the unreasonable one
    persists in trying to adapt the world to himself. Therefore all
    progress depends on the unreasonable man. -- George Bernard Shaw


    --- PyGate Linux v1.5.17
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From Charles Curley@3:633/10 to All on Friday, June 26, 2026 00:50:01
    On Mon, 22 Jun 2026 15:40:11 -0700
    David Christensen <dpchrist@holgerdanske.com> wrote:

    It looks like there is a newer kernel for Trixie. It is best to file
    bug reports against current packages. Can you test it?

    https://packages.debian.org/stable/kernel/linux-image-6.12.94+deb13-amd64

    Tested. It came up with all sorts of fails.

    charles@hawk:~$ uname -a
    Linux hawk 6.12.94+deb13-amd64 #1 SMP PREEMPT_DYNAMIC Debian 6.12.94-1 (2026-06-20) x86_64 GNU/Linux
    charles@hawk:~$

    I gather you are suggesting I file a bug against this kernel, linux-image-6.12.94+deb13-amd64 6.12.94-1.

    --
    Does anybody read signatures any more?

    https://charlescurley.com
    https://charlescurley.com/blog/

    --- PyGate Linux v1.5.18
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From David Christensen@3:633/10 to All on Friday, June 26, 2026 03:10:02
    On 6/25/26 15:40, Charles Curley wrote:
    On Mon, 22 Jun 2026 15:40:11 -0700
    David Christensen <dpchrist@holgerdanske.com> wrote:

    It looks like there is a newer kernel for Trixie. It is best to file
    bug reports against current packages. Can you test it?

    https://packages.debian.org/stable/kernel/linux-image-6.12.94+deb13-amd64

    Tested. It came up with all sorts of fails.

    charles@hawk:~$ uname -a
    Linux hawk 6.12.94+deb13-amd64 #1 SMP PREEMPT_DYNAMIC Debian 6.12.94-1 (2026-06-20) x86_64 GNU/Linux
    charles@hawk:~$

    I gather you are suggesting I file a bug against this kernel, linux-image-6.12.94+deb13-amd64 6.12.94-1.


    That would make sense, especially since Linux 6.12 appears to be the
    kernel for Debian Stable (Trixie):

    https://packages.debian.org/trixie/all/allpackages


    David

    --- PyGate Linux v1.5.18
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From Charles Curley@3:633/10 to All on Friday, June 26, 2026 05:50:02
    On Thu, 25 Jun 2026 18:07:21 -0700
    David Christensen <dpchrist@holgerdanske.com> wrote:

    That would make sense, especially since Linux 6.12 appears to be the
    kernel for Debian Stable (Trixie):

    https://packages.debian.org/trixie/all/allpackages

    OK, will do.

    However, things are back up in the air. I rebooted to 7.0.10, and ran
    some backups. The prior testing has been all reading: checksum
    verification. Nothing on the disk actually changed. This backs from the
    SSD to the RAID array. I had several directories fail with the error
    message "failed: Bad message (74)", whatever that means.

    I immediately fscked the logical volume. No errors. I then diffed the directories against the originals. There were for instances of files
    found in the backups and not in the originals. Otherwise the "failed" directories were intact and duplicated the originals. I conjecture that
    rsync attempted to delete them and failed to do so. I conjecture that
    the next pass will get them.

    --
    Does anybody read signatures any more?

    https://charlescurley.com
    https://charlescurley.com/blog/

    --- PyGate Linux v1.5.18
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)
  • From David Christensen@3:633/10 to All on Saturday, June 27, 2026 01:50:02
    On 6/25/26 20:44, Charles Curley wrote:
    On Thu, 25 Jun 2026 18:07:21 -0700 David Christensen wrote:
    That would make sense, especially since Linux 6.12 appears to be the
    kernel for Debian Stable (Trixie):

    https://packages.debian.org/trixie/all/allpackages

    OK, will do.


    Thank you.


    However, things are back up in the air. I rebooted to 7.0.10, and ran
    some backups. The prior testing has been all reading: checksum
    verification. Nothing on the disk actually changed. This backs from the
    SSD to the RAID array. I had several directories fail with the error
    message "failed: Bad message (74)", whatever that means.

    I immediately fscked the logical volume. No errors. I then diffed the directories against the originals. There were for instances of files
    found in the backups and not in the originals. Otherwise the "failed" directories were intact and duplicated the originals. I conjecture that
    rsync attempted to delete them and failed to do so. I conjecture that
    the next pass will get them.


    Are you booting the OS disk and running backups, or are you booting live
    media and backing up?


    Please post a console session that shows prompts, commands entered, and
    output displayed. If you are using shell scripts, please enable the
    "xtrace" option.


    David

    --- PyGate Linux v1.5.18
    * Origin: Dragon's Lair, PyGate NNTP<>Fido Gate (3:633/10)