https://gitlab.synchro.net/main/sbbs/-/commit/ce3b73460e452f53f639b8e6
Modified Files:
ctrl/text.dat docs/v322_new.md exec/load/text.js src/sbbs3/bat_xfer.cpp filedat.c filedat.h main.cpp sbbs.h text.h text_defaults.c text_id.c upload.cpp src/smblib/smbfile.c smblib.h
Log Message:
Batch uploads: read each directory's file index once per batch (#1132)
The duplicate-file scan that follows every upload opened and read every dupe-checked directory's file index (the .sid) for every uploaded file,
so a batch of N files cost N passes over every index: on a system with
1,300 directories that is 1,300 opens and a full index read per file,
and each open is a round trip when the data lives on a file server.
Add a duplicate-check cache (dupe_cache_* in filedat.c) holding the size
and hash values of each directory's files, read from its index the first
time the directory is consulted and kept for the life of the batch. process_batch_upload_queue() creates it for the whole batch, uploadfile() consults it when present and records each file it adds so a later file in
the same batch is checked against it too. A single-file upload is
unchanged: the same one pass over the indexes it always made.
smb_findfile()'s size-and-hash comparison moves into smb_filehash_match()
so the cache and the index search apply the same rule.
Also print the file's name before its upload testers, hashing and
duplicate scan begin (new ProcessingUploadedFile text string), so the
user can tell which file a long post-processing step belongs to.
Verified on a scratch install with a fake batch protocol: within one
batch, a file with the same content as an earlier one was reported as
already uploaded, and so was one in a later batch; strace showed each directory's index opened once per batch by the duplicate scan.
Co-Authored-By: Claude Fable 5.1 <
noreply@anthropic.com>
---
þ Synchronet þ Vertrauen þ Home of Synchronet þ [vert/cvs/bbs].synchro.net