Repository navigation
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #707 +/- ##
==========================================
+ Coverage 93.87% 94.36% +0.48%
==========================================
Files 59 59
Lines 3218 3227 +9
==========================================
+ Hits 3021 3045 +24
+ Misses 197 182 -15
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR reads images in small pieces and runs uploads in worker threads. The request thread can then handle other work while an upload runs.
For 16 concurrent 1 MiB images, the batch took 84 ms instead of 435 ms in the local boto3 benchmark. That is about 81% less time, or 5.2x faster than baseline. The test uses the real boto3 transfer code with local simulated storage calls. It does not send data to a real S3 server.
flowchart TB I["Image file"] --> A["Before: read the whole image<br/>for each hash and the type check"] A --> B["Upload on the request thread<br/>Other requests must wait for that work"] I --> C["After: read 64 KiB at a time<br/>Calculate both hashes in one pass"] C --> D["Upload in a worker thread<br/>At most 8 uploads per API process"] D --> E["The request thread can<br/>handle other work"]A hash is a checksum of the file's bytes. This PR keeps the same SHA256 and MD5 checksums. It reads an 8 KiB header to identify the file type. It releases each hashing piece before reading the next one.
The two smaller-image batches are 1.3–1.5x faster than the earlier PR version. Peak Python allocations for the 1 MiB batch fell from 1.10 to 0.77 MiB versus baseline. For the 4 MiB batch, they fell from 4.21 to 1.21 MiB.
For one 64 MiB image with simulated storage, peak total process memory fell from 177.7 to 114.5 MiB, about 36% less. That test does not use boto3's real multipart upload path. The boto3 tests with smaller images showed about the same total process memory after warmup.
There is a tradeoff. Eight workers complete the measured batches sooner than four workers, but use more memory. Four workers with the same small hashing pieces took 127/209 ms for the two boto3 batches. For one 1 MiB image with no simulated storage delay, the worker overhead increased time from 3.25 to 4.53 ms.
Benchmark method, implementation, and checks
Python 3.11.15. Baseline:
25f3d2ed10bd. Earlier PR version:e0783ad4(256 KiB hashing pieces; four workers). Each condition uses a fresh process and seven timed calls after warmup. Timing excludes imports and input construction. Python allocation tracing runs separately. The consuming-stub cases run in random order; the boto3 cases run in a fixed order. Results measure batch completion, not individual request latency or full API load.The boto3 tests use actual S3Transfer scheduling and buffering. Local put-object and metadata calls each wait 10 ms, then consume the bytes or return metadata. No HTTP traffic occurs. A wrapper prevents boto3 from closing fixture files between samples. It does not change transfer buffering or scheduling. The 64 MiB cases use a byte-consuming storage stub instead of the SDK transfer.
The request-loop delay fell from about 522–539 ms to 7–10 ms in the boto3 cases. Total process memory was within 1 MiB of the earlier PR version. Traced Python allocations exclude native memory and are not total process memory.
Use AnyIO workers for hashing, upload, and verification. Limit concurrent uploads to eight per API process; boto3 can use additional transfer threads. Keep cancellation shielding until the worker finishes with the file. Keep file keys, prefixes, MD5/ETag verification, and cleanup after corruption. Enable thread tracing in both root and container coverage configurations.
Four tests cover bounded reads, exact bytes and file keys, upload/corruption errors, cleanup, worker limits, and request-loop responsiveness. Eleven external comparisons cover spooled files, file types, prefixes, and completion after cancellation. CI passes: 670 backend tests passed with one skipped, plus two separate Telegram checks. Lint, formatting, type checks, and coverage checks pass; patch coverage is 100%.
Net diff: +12 production lines; +90 including tests. The coverage-setting replacements add no lines. No dependency changes or benchmark files in the diff. Real network latency and multipart buffering can change results. The existing multipart ETag check is unchanged; the 64 MiB stub tests do not prove successful real multipart uploads. #664 has no storage changes.