summaryrefslogtreecommitdiff
AgeCommit message (Collapse)AuthorFilesLines
2026-08-20pack-objects: trace pack bytes writtenFriel2-0/+29
We want to measure how compression settings affect push performance on the client. Different settings can produce different-sized packs from the same objects. Trace2 records the object count, but we also need the pack size to compare those settings. Add a write_pack_file/wrote_bytes Trace2 datum alongside write_pack_file/wrote. Count packs written to stdout or disk, including each pack's header and trailing checksum. When pack.packSizeLimit splits the output, report the sum of the pack sizes. Signed-off-by: Friel <friel@openai.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-20The 16th batchJunio C Hamano1-0/+16
Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-20Merge branch 'kh/doc-trailers'Junio C Hamano1-24/+64
Documentation for 'git interpret-trailers' has been updated to explain the format of trailer keys (alphanumeric characters and hyphens), replace outdated terminology, define key terms upfront, and document how comment lines in the input are treated. * kh/doc-trailers: doc: interpret-trailers: document comment line treatment doc: interpret-trailers: rewrite new-trailers paragraphs doc: interpret-trailers: commit to “trailer block” term doc: interpret-trailers: join new-trailers again doc: interpret-trailers: add key format example doc: interpret-trailers: explain key format doc: interpret-trailers: explain the format after the intro doc: interpret-trailers: not just for commit messages doc: interpret-trailers: use “metadata” in Name as well doc: interpret-trailers: replace “lines” with “metadata” doc: interpret-trailers: stop fixating on RFC 822
2026-08-20Merge branch 'ps/odb-make-creation-pluggable'Junio C Hamano17-58/+168
The creation of the on-disk data structures for the object database has been made pluggable, allowing future backends to customize their setup. As part of this, the initialization of the object database has been deferred, and the loading of the loose-object map has been detangled from repository initialization. * ps/odb-make-creation-pluggable: odb: make creation of on-disk structures pluggable odb/source: introduce function to map source type to name setup: defer object database creation setup: handle ODB-related environment variables in `odb_new()` setup: detangle loading of loose object maps loose: load loose object map for the correct source
2026-08-20Merge branch 'hn/branch-delete-merged'Junio C Hamano6-31/+841
The 'git branch' command has been taught the '--delete-merged' option to remove local branches that are already merged into their tracked remote-tracking branches. * hn/branch-delete-merged: branch: add --dry-run for --delete-merged branch: add branch.<name>.deleteMerged opt-out branch: add --delete-merged <pattern> branch: prepare delete_branches for a bulk caller branch: let delete_branches skip unmerged branches on bulk refusal branch: convert delete_branches() to a flags argument branch: add --forked filter for --list mode
2026-08-19odb: handle `OBJECT_INFO_DIE_IF_CORRUPT` genericallyPatrick Steinhardt5-40/+52
When a lookup with `OBJECT_INFO_DIE_IF_CORRUPT` fails we want to die in case the object exists, but cannot be read. This flag is handled in two different spots right now: - `do_oid_object_info_extended()` calls `has_packed_and_bad()` to check whether the object is known to be corrupt in any packfile. This function reaches into the internals of the packed source and thus breaks the abstraction provided by our object sources. - The loose source handles the flag itself and dies directly in `read_object_info_from_path()`, which means that we die even in cases where another source may still have a good copy of the object. Besides being inconsistent, it also ties us to the specific backend used by the database sources because `has_packed_and_bad()` assumes that they use the "files" backend. Any other backend will instead cause us to die when calling `odb_source_files_downcast()`, even if the object was simply nonexistent. In the preceding commits we've carved out the infrastructure to make this mechanism fully generic. On the one hand, all backends now tell us whether the object is missing or corrupt via their return values. And on the other hand, they have been taught to provide a readable error message to the caller. Adapt `do_oid_object_info_extended()` to use those new mechanisms. This means that we won't die immediately anymore when a loose object is corrupt, and we properly handle backends other than the "files" backend. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-19odb/source: allow `read_object_info()` to bubble up error messagesPatrick Steinhardt9-33/+73
When reading an object fails even though it exists, the sources know best what exactly went wrong and where the corrupt object is located. This information is lost though when bubbling up the error to the object database layer, which forces that layer to reconstruct it after the fact. This is exactly what `do_oid_object_info_extended()` does via `has_packed_and_bad()`, but that function only really knows to handle the "files" backend by reaching into its internals. Introduce a new `errmsg` parameter for the `read_object_info()` callback that sources are expected to populate with a human-readable message in case reading the object has failed. Adapt the packed and loose sources to populate the buffer with the messages that we ultimately want to surface to the user. For now, all callers are adapted to pass a `NULL` pointer. We will add a user of this new infrastructure in a subsequent commit. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-19odb/source: let callers discern missing and corrupt objectsPatrick Steinhardt6-18/+42
As explained in the preceding commits, reading objects can either fail because the object truly does not exist or because it exists, but its data is corrupt. Some callers do care about this distinction, but there is no way to tell these two cases apart right now. Introduce a new `ODB_READ_NOT_FOUND` value that ought to be returned by the backends in case the object truly does not exist and adapt backends to use it. Note that we don't yet return this error from `odb_read_object_info()` itself. This will be fixed in a subsequent commit. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-19odb/source: introduce error status when reading objectsPatrick Steinhardt7-39/+46
The `read_object_info()` callback of `struct odb_source` is documented to return a negative error code in case reading the object has failed, and zero otherwise. This is overly broad though, as there are two very different kinds of failures: - The object may not exist in the source at all. - The object exists, but reading it has failed, for example because its on-disk state is corrupt. This distinction matters to callers: when an object is corrupt in one source we may still find a good copy of it in another source, so we may still be able to proceed with a given operation. The "packed" source already distinguishes these cases by returning a positive value for missing objects and a negative value in case reading the object has failed. But it is the only such source that distinguishes those cases, and the returned value is translated into a negative error code by the "files" backend anyway. Introduce a new error status that is specific to reading objects and adapt the infrastructure to return it. For now, we only discern successful reads from generic failures, which mostly matches the status quo. In subsequent commits though we're about to add an error that explicitly tells the caller that an object does not exist. Note that we keep the "packed" backend as-is with its positive return code for missing objects. This will be fixed in the next commit. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-19odb/source-packed: flag known-bad objects as corrupt and not missingPatrick Steinhardt7-16/+36
When reading packed objects we know to tell apart missing objects and corrupt objects by returning a positive error code in the former case, and a negative one in the latter case. We do that by distinguishing between errors returned by `find_pack_entry()`, which yields the offset of the object, and `packed_object_info()`, which reads the object contents. But even though we already distinguish those cases when reading packed objects, the logic is broken in case a caller tries to read an object that has been marked as corrupt. In that case, `find_pack_entry()` will tell us that the object in question does not exist, and consequently we'll not flag the object as corrupt but as missing. Fix this issue by bubbling up whether the object is corrupt and, if so, which packfile contains the corrupted object. Note that we don't yet need the information about the specific packfile, so we could've just as well made this a `bool *corrupted` pointer. But we'll need information about the containing packfile in a subsequent commit so that we can generate a proper error message telling the user which packfile contains the broken object. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-19completion: zsh: support completion after "git -C <path>"Lutz Lengemann1-1/+24
The zsh completion wrapper does not handle the global -C option, so git -C <path> <command> <TAB> offers nothing. -C is not part of the _arguments specification, and the wrapper hard-codes __git_cmd_idx=1, i.e. it assumes that the command is the first argument, so the bash helpers look at the wrong word. The latter is not specific to -C; the assumption breaks after any global option, e.g. "git -p checkout <TAB>" does not complete branch names. Add -C to the specification, and find the command by skipping over the global options and, where they take one, their arguments, as __git_main in git-completion.bash does. The index is one less than zsh's, as the helpers count the words from zero. Collect the paths given to -C into __git_C_args, or else the helpers run git in the current directory and fail to resolve the aliases and refs of the repository the command runs in. The argument of a -C is still completed without regard for the -C options before it, i.e. "git -C dir -C <TAB>" offers the directories in ".", not the ones in "dir". Signed-off-by: Lutz Lengemann <lutz@lengemann.net> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-18The 15th batchJunio C Hamano1-0/+16
Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-18Merge branch 'hn/bisect-reset-when-found'Junio C Hamano4-14/+285
The 'git bisect' command has been taught a '--reset-when-found[=<where>]' option that tells the command to automatically run 'git bisect reset' to jump back to the original state or to the found culprit. * hn/bisect-reset-when-found: bisect: add --reset-when-found to leave when done bisect: let bisect_reset() optionally check out quietly
2026-08-18Merge branch 'ps/writev'Junio C Hamano12-7/+190
A compatibility wrapper for writev(3p) has been reintroduced, including fixes for CMake build and 'MAX_IO_SIZE' limits on NonStop. Calls to write(3p) in send_sideband() and cat_blob() have been refactored to use writev(3p) wrappers to reduce syscall overhead. * ps/writev: fast-import: use writev(3p) to send cat-blob responses sideband: use writev(3p) to send pktlines wrapper: properly handle MAX_IO_SIZE in writev(3p) wrapper: introduce writev(3p) wrappers compat/posix: introduce writev(3p) wrapper
2026-08-18Merge branch 'kh/doc-refs-migrate-limitations'Junio C Hamano1-15/+15
The known limitations of the ref format migration in 'git refs' have been moved to be displayed as a warning admonition directly under the description of the 'migrate' subcommand, improving visibility. A reference to 'git-maintenance' has also been corrected to use the 'linkgit' macro. * kh/doc-refs-migrate-limitations: doc: refs: linkgit to git-maintenance(1) doc: refs: put ref migration warning under the command
2026-08-17doc: format-rev: use [synopsis] on code blockKristoffer Haugsbakk1-2/+3
This code block uses the placeholder `<subject>`. Let’s highlight this placeholder properly by using the `synopsis` open block definition which was introduced in a34d1d53 (doc: convert git-show to synopsis style, 2026-02-06). This renders the block like a code block but with emphasis styling on placeholders, just like inline-verbatim (`) in running text. Yes, note that open blocks since commit a34d1d53 can, on synopsis-style docs like this one, be immediately preceded by `[synopsis]`, just like the command synopsis is: [synopsis] (EXPERIMENTAL!) git format-rev - [...] Cf. verse-style: [verse] 'git name-rev' [...] Signed-off-by: Kristoffer Haugsbakk <code@khaugsbakk.name> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-17doc: format-rev: quote subject placeholder before and afterKristoffer Haugsbakk1-2/+2
We first talk about just `%s`, but then show the result with quotes. That is inconsistent. Let’s use quotes both in the format as well as in the result. The implied input here, which is not spelled out for brevity, is: Did we not fix this in <commit object name>? Which is then supposed to be formatted to `"<subject>"`. Signed-off-by: Kristoffer Haugsbakk <code@khaugsbakk.name> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-17odb: drop `alternates_db` fieldPatrick Steinhardt2-15/+9
The `struct object_database::alternates_db` field tracks the value of the "GIT_ALTERNATE_OBJECT_DIRECTORIES" environment variable and is used in `odb_prepare_alternates()`. It's not necessary to store it as a separate field anymore though, as we stopped lazy-loading alternates. Consequently, we can simply pass it to `odb_prepare_alternates()` via `odb_new()` now. Do so and remove the field. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-17odb: drop `loaded_alternates` fieldPatrick Steinhardt2-10/+1
The `struct object_database::loaded_alternates` field tells us whether or not alternates have been loaded already. This field was useful before the preceding commit as we were indeed lazy-loading alternates. But now that we started to eagerly load them we can assume them to be loaded after `odb_new()`, and hence the field does not serve any purpose anymore. Remove it. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-17odb: eagerly initialize alternatesPatrick Steinhardt11-46/+4
When creating the object database we initialize the main object database source, but we don't yet initialize its alternates. Instead, we have many calls to `odb_prepare_alternates()` cluttered around the code base whenever we are about to iterate through the sources. This lazy loading doesn't really add much value: the moment where we read any object we _have_ to load the alternates anyway. So given that most of our commands would access the object database this optimization is not really buying us much in the first place. Quite on the contrary, it makes the code harder to understand and is a potential source of bugs in case any callsite forgot to prepare alternates before we iterate through the sources. Historically though there was a reason why we deferred lazy-loading: it may happen that the repository has "core.ignoreCase" configured, and we use that to deduplicate the list of alternates in case we had the same alternate configured multiple times, but with different casing. We used to initialize the object database before we had fully configured the owning repository though, and consequently we couldn't access that configuration yet. This has changed in the preceding commit though where we started to parse "core.ignoreCase" manually. Eagerly prepare alternates both when creating the object database and when flushing its caches. Drop the now-unneeded calls to prepare the alternates that are scattered across the code base. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-17odb: decouple source path comparisons from `the_repository`Patrick Steinhardt3-22/+78
When registering alternates we deduplicate object database sources by their path so that the same source won't be added twice. Ever since cf2dc1c238 (speed up alt_odb_usable() with many alternates, 2021-07-07) this duplicate check is backed by a map keyed by the source's path, using `fspathhash()` and `fspatheq()` as hash and equality functions, respectively. These functions are problematic in this context for two reasons: - They implicitly depend on `the_repository` instead of the repository that owns the object database. - They derive case-sensitivity from `repo_ignore_case()`, which returns a default value in case the repository's configuration has not been parsed yet. Object database sources may be registered before that is the case, so the answer may flip depending on when a source gets registered. Fix this by making the comparison self-contained in the object database. Instead of using `fspathhash()` and `fspatheq()` we resolve "core.ignoreCase" manually and then use the correct comparison function based on the result. This requires us to migrate to a `struct hashmap`, as the khash interface does not give us the ability to pass an arbitrary payload to these functions, and hence we'd have to use global state to decide which of those to use. Note that we can unconditionally use `strihash()` to compute entry hashes regardless of case sensitivity: a hash function only needs to guarantee that equal keys have equal hashes, and a case-insensitive hash satisfies this requirement for both case-sensitive and case-insensitive equality. Overall it's quite debatable whether all of this complexity really is worth it, out of two reasons: - We could linearly search through all sources to find duplicates. But the mentioned commit cares about cases with thousands of alternates, and a linear search would of course regress performance quite a bit. This doesn't really feel like a reasonable case to care about, but I don't feel comfortable regressing it anyway. - It's dubious whether we should handle "core.ignoreCase" in the first place. The downside would be that we might add the same alternate multiple times with different casing. But this is an edge case, and it's not even fully fixed because we don't resolve symlinks or mountpoints, either. So for now, keep this infrastructure in-place while removing the global dependency on `the_repository`. We may want to revisit this in the future though. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-17setup: create ref and object databases after config is writtenPatrick Steinhardt1-6/+6
When creating a new repository we create both the reference and object databases after we have finalized the repository. This ensures that those subsystems find a fully-configured repository at the time where they are asked to create their own on-disk data structures. There is one exception though: while we have already fully configured the repository at this point, we haven't yet written both "core.sharedRepository" and "receive.denyNonFastforwards". The latter configuration doesn't really matter to us, but the first one does as the "files" object database source reads it. This doesn't cause any problems right now, but it will in a subsequent patch where we will start to read "core.ignoreCase" when creating the object database. Move the initialization of both of these data structures towards the end of `init_db()`. The only thing that now comes after is status reporting, but that's it. Signed-off-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-17object-name: avoid use-after-free in get_oid_with_context_1()Shlok Kulshreshtha2-6/+19
When a ":<path>" argument names a relative path, resolve_relative_path() returns a newly allocated string and "cp" is pointed at it: new_path = resolve_relative_path(repo, cp); if (!new_path) { namelen = namelen - (cp - name); } else { cp = new_path; namelen = strlen(cp); } From there on "cp" and "new_path" name the same allocation. Later the memory location that "new_path" points to is freed. free(new_path); if (reject_tree_in_index(repo, only_to_die, ce, stage, prefix, cp)) But here the reject_tree_in_index() passes "cp" to diagnose_invalid_index_path(), which calls strlen() on it, looks it up in the index, and formats it into its messages, allocating as it goes. All of this reads memory that has already been freed. Collapse the two exits into one to ensure a single free() that happens after the last use. Three things have to coincide to reach this: 1. The path has to be relative, or nothing is allocated and "cp" still points into the argument. 2. The entry found has to be a sparse directory, which needs a sparse index. 3. The argument has to get past the check in die_verify_filename() that skips a leading ':' followed by a non-alphanumeric, so ":0:./dir/" arrives here where ":./dir/" does not. Add a test to t1092 that covers the combination. It fails under SANITIZE=address without the change to object-name.c. This was reported in [1], and the shape used here was suggested in review [2], but that series was not rerolled and the fix never landed. [1] https://lore.kernel.org/git/cf6bcdb43e5b4abab464c30a914d64dc8e7a9925.1655336146.git.gitgitgadget@gmail.com/ [2] https://lore.kernel.org/git/xmqqy1xxw7rc.fsf@gitster.g/ Reported-by: Johannes Schindelin <Johannes.Schindelin@gmx.de> Original-patch-by: Johannes Schindelin <Johannes.Schindelin@gmx.de> Helped-by: Junio C Hamano <gitster@pobox.com> Suggested-by: Patrick Steinhardt <ps@pks.im> Signed-off-by: Shlok Kulshreshtha <diy2903@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-15The 14th batchJunio C Hamano1-0/+21
Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-15Merge branch 'jc/add-resolved'Junio C Hamano10-99/+321
'git add' has been taught a new '--resolved' option to stage conflict-resolved paths, while leaving unrelated local changes unstaged. It scans the unmerged paths for leftover conflict markers and aborts if any are found. * jc/add-resolved: add: introduce '--resolved' option read-cache: add remove_file_from_index_with_flags() merge-ll: consolidate conflict marker scanning logic read-cache: reindent
2026-08-15Merge branch 'kl/t7528-ssh-agent-for-csh-users'Junio C Hamano1-1/+1
The 'ssh-agent' tests in 't7528' have been fixed to work when the user's login shell is csh-like, by explicitly passing '-s' to 'ssh-agent' to force Bourne shell syntax. * kl/t7528-ssh-agent-for-csh-users: t7528: fix failure under csh
2026-08-15Merge branch 'tn/packfile-uri-concurrency'Junio C Hamano8-46/+305
Concurrent downloads of packfiles via packfile URIs and dumb HTTP are safer by avoiding concurrent appends to the staging file. Opening in read-write mode with separate file offsets prevents corruption and preserves resumability. 'fetch-pack' now tolerates pre-existing '.keep' files. * tn/packfile-uri-concurrency: fetch-pack: accept "pack" output for packfile URIs http: permit unlinking partial packs on Windows http: avoid concurrent appends to partial packs http: accept HTTP 416 for complete partial packs http: avoid closing index-pack input twice http-fetch: correct --index-pack-arg documentation
2026-08-15Merge branch 'jm/t0213-skip-emulated-ancestry-tests'Junio C Hamano1-3/+6
The 'TRACE2_ANCESTRY' prerequisite in the 't0213' test script has been refined to avoid failures under user-mode emulation by verifying that the ancestry collector reports the expected process names rather than the emulator binary name. * jm/t0213-skip-emulated-ancestry-tests: t0213: skip ancestry tests under user-mode emulation
2026-08-15repository: move fetch_if_missing into struct repositoryTian Yuchen15-49/+51
The global variable 'fetch_if_missing' controls whether a missing object check should prompt a lazy fetch from a promisor remote. In order to continue the libification effort, move it into 'struct repository' and initialize it to 1 by default to keep the previous behavior. builtin/fetch-pack.c, builtin/fsck.c, and builtin/rev-list.c are entered via commands marked RUN_SETUP in git.c:commands[]. Their 'repo' parameter is only NULL when '-h' is given outside of a repository, in which case either show_usage_if_asked() or parse_options()'s own '-h' handling exits the process before returning. We can therefore drop their UNUSED markers and assign to 'repo' directly. builtin/index-pack.c is entered via RUN_SETUP_GENTLY, so its 'repo' pointer can be NULL any time it is run outside of a repository, not only with '-h'. We keep a NULL check there and fall back to 'the_repository'. builtin/pack-objects.c needs two adjustments to make 'repo' reach every 'fetch_if_missing' call site: 'read_stdin_packs()' now takes a 'struct repository *'; 'option_parse_missing_action()', which is registered as an OPT_CALLBACK, receives a 'repo' through the option's 'value' field now. Additionally, update the partial clone documentation to reflect that this is now a per-repository flag. Mentored-by: Christian Couder <christian.couder@gmail.com> Mentored-by: Ayush Chandekar <ayu.chandekar@gmail.com> Mentored-by: Olamide Caleb Bello <belkid98@gmail.com> Signed-off-by: Tian Yuchen <cat@malon.dev> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-14chdir-notify.h: Removed unused param 'name'Colin Hinton10-44/+30
The `name` parameter in `chdir_notify_entry` was only ever used by chdir_notify_reparent() to produce trace output. That function was removed in 5bf546755c (chdir-notify: drop unused `chdir_notify_reparent()`, 2026-06-25), which left `name` with no remaining consumers. Prior to that removal, most callers had already stopped passing a meaningful name, switching to NULL in 1f43ff2c7e (refs: unregister reference stores from "chdir_notify", 2026-06-25) and 0de2467e6c (odb/source-packed: start converting to a proper `struct odb_source`, 2026-06-17). Since no caller has populated `name` with real data for some time, and its last consumer is gone, drop it from chdir_notify_register(), chdir_notify_unregister(), and the callback signature to simplify the API. Signed-off-by: Colin Hinton <colinlewishinton@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-14doc: fix typo in submitting patchesSwapnil Saste | INDIA1-1/+1
Remove the article "an" before "incremental updates". Signed-off-by: Swapnil Saste | INDIA <theswapnilsaste@gmail.Com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13diff: avoid misleading statement about -l optionElijah Newren1-1/+1
In commit 6623a528e00b (doc: clarify documentation for rename/copy limits, 2021-07-15), the wording around rename limit options and config variables were updated to point out that only the quadratic portion of rename detection (or "exhaustive portion of rename/copy detection" as used in that commit) was limited by these options, because exact rename detection and basename-guided rename detection (which both run in time linear in the number of files) still run before this limit is checked. However, the short help message wasn't updated at the time; update it too. Signed-off-by: Elijah Newren <newren@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13builtin/repack: add guards for --drop-filteredSiddharth Shrimali3-0/+96
--drop-filtered removes local promisor blobs. That is only safe when the repository is not mid-operation and when the blobs are not actively in use, so add two guards, both skipped for bare repositories which have neither a worktree nor an index. First, refuse to run while a merge, rebase, am, cherry-pick, revert, or bisect is in progress. During these operations the working tree and index are in an intermediate state, and rewriting packs and deleting objects underneath a half-finished operation is unsafe. Second, refuse to drop a blob that the current index references. Such a blob is needed by the working tree, so dropping it would only cause the next command that touches the worktree to lazy-fetch it straight back, reclaiming nothing. The offending path is reported so the user can see why the drop was refused. Mentored-by: Christian Couder <christian.couder@gmail.com> Mentored-by: Siddharth Asthana <siddharthasthana31@gmail.com> Signed-off-by: Siddharth Shrimali <r.siddharth.shrimali@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13builtin/repack: actually drop filtered promisor blobsSiddharth Shrimali3-4/+50
Make --drop-filtered remove the enumerated promisor blobs instead of only listing them. The drop set is computed before repack_promisor_objects() runs, and on a real run it is passed in so the rebuilt promisor pack omits those blobs. --drop-filtered implies -d so the old promisor packs, which still contain the dropped blobs, are removed. Without this the blobs would survive in the redundant packs. The existing repack machinery performs the write-before-delete and fsync, so the drop is crash-safe. The dropped blobs become absent locally but remain recoverable from the promisor remote, so a later access lazy-fetches them back transparently. --dry-run keeps its previous behavior, i.e. it lists the candidates and changes nothing. Mentored-by: Christian Couder <christian.couder@gmail.com> Mentored-by: Siddharth Asthana <siddharthasthana31@gmail.com> Signed-off-by: Siddharth Shrimali <r.siddharth.shrimali@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13builtin/repack: enumerate promisor blobs for --drop-filteredSiddharth Shrimali4-2/+189
Add enumeration logic for --drop-filtered. In --dry-run mode, print the OIDs of locally-held promisor blobs that exceed the filter threshold, as candidates for removal. Reading from write_filtered_pack() cannot work for partial clones. git repack routes promisor objects through a separate path: repack_promisor_objects() repacks them first, and the main pack-objects run uses --exclude-promisor-objects. By the time write_filtered_pack() runs, the promisor blobs are already consumed by the main pack. The filtered pack is always empty on a partial clone. Instead, walk promisor objects directly via odb_for_each_object() with ODB_FOR_EACH_OBJECT_PROMISOR_ONLY, collecting all promisor blobs into an oidset. The blobs exceeding the filter threshold are then selected using list_objects_filter__filter_oidset(). Every object enumerated this way is a promisor object, so it is recoverable from the promisor remote in the same sense as the rest of a partial clone, as long as the remote still has it. This holds without a separate is_promisor_object() check. A future implementation can verify availability against the remote directly once a client-side remote-object-info query exists. OBJECT_INFO_SKIP_FETCH_OBJECT is passed to every object info query so enumeration never triggers a lazy fetch. The enumeration collects candidates into a caller-provided oidset and --dry-run prints them. Actually removing the objects, together with the required promisor-remote verification, is written in a later commit. Mentored-by: Christian Couder <christian.couder@gmail.com> Mentored-by: Siddharth Asthana <siddharthasthana31@gmail.com> Signed-off-by: Siddharth Shrimali <r.siddharth.shrimali@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13repack-promisor: allow excluding objects from the rebuilt promisor packSiddharth Shrimali3-3/+18
Add a to_drop oidset parameter to repack_promisor_objects(). When it is non-NULL, write_oid() omits those objects from the rebuilt promisor pack. This is the mechanism --drop-filtered will use to remove promisor blobs, i.e. rebuild the promisor pack without them. All existing callers pass NULL, so behavior is unchanged. Mentored-by: Christian Couder <christian.couder@gmail.com> Mentored-by: Siddharth Asthana <siddharthasthana31@gmail.com> Signed-off-by: Siddharth Shrimali <r.siddharth.shrimali@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13list-objects-filter: add list_objects_filter__filter_oidset()Siddharth Shrimali2-0/+61
The existing filter entry point, list_objects_filter__filter_object(), is built around the object-walk path: it expects traversal context and provisional omit sets, and is meant to be called as objects are visited during a walk. A caller that already has a set of OIDs in hand and only wants to know which ones a filter would select has no usable entry point into the filter API. --drop-filtered is exactly such a caller: it collects promisor blobs into an oidset and needs to know which of them exceed the filter threshold, without performing an object walk. Add a helper, list_objects_filter__filter_oidset(), that takes a set of OIDs and populates an "omitted" set with those that would be filtered out by the given filter options. Only blob:limit=N filters are supported for now. This helper does not actually reuse the existing filter machinery. It reimplements the blob:limit size check directly. That machinery is tied to the object-walk path and cannot easily be driven from a plain oidset. A NEEDSWORK comment marks this so the helper can later be refactored to reuse the real filter logic instead of duplicating it. OBJECT_INFO_SKIP_FETCH_OBJECT is passed when reading object info so the helper never triggers a lazy fetch. Mentored-by: Christian Couder <christian.couder@gmail.com> Mentored-by: Siddharth Asthana <siddharthasthana31@gmail.com> Signed-off-by: Siddharth Shrimali <r.siddharth.shrimali@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13builtin/repack: add --drop-filtered and --dry-run optionsSiddharth Shrimali3-3/+127
Add two new command-line options to 'git-repack': --drop-filtered: intended to eventually delete objects that match the filter specification. Requires --filter and -a, and is incompatible with --filter-to. --dry-run: show which objects would be dropped without making any changes. Only meaningful with --drop-filtered. Keep --dry-run as a separate option rather than folding it into --drop-filtered (e.g. --drop-filtered=dry-run), to stay consistent with the --dry-run option other Git commands already provide and to leave room for it to describe other repack behavior later. A --drop-filtered=<mode> form can still be added later if more drop-specific modes are needed. --drop-filtered also requires a promisor remote to be configured, since dropping objects without a remote to fetch them back from would be permanent data loss. --drop-filtered is incompatible with bitmap writing: filtering breaks the "all objects in one pack" closure that bitmaps require. Detect an explicit -b/--write-bitmap-index on the command line with a dedicated option callback that sets a "write_bitmaps_given" flag, so it can be distinguished from a repack.writeBitmaps configuration value even when config already enables bitmaps. An explicit -b is reported as a conflict, while a config-provided default is silently disabled for the duration of the command. These options currently only perform validation. The actual enumeration and deletion will be added in follow-up commits. Mentored-by: Christian Couder <christian.couder@gmail.com> Mentored-by: Siddharth Asthana <siddharthasthana31@gmail.com> Signed-off-by: Siddharth Shrimali <r.siddharth.shrimali@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13completion: complete 'git history split' pathspecsVincent Mailhol2-0/+19
Arguments following the required revision of "git history split" are pathspecs. Complete them from tracked paths, including after an explicit "--". Signed-off-by: Vincent Mailhol <mailhol@kernel.org> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13completion: complete 'git history --update-refs' valuesVincent Mailhol2-1/+9
The "--update-refs" option accepts either "branches" or "head". Complete these values for the documented --update-refs=<value> form. While parse-options also accepts the split --update-refs <value> form, it is not documented. Omit it from completion as a trade-off for code simplicity. Signed-off-by: Vincent Mailhol <mailhol@kernel.org> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13completion: complete 'git history --empty' valuesVincent Mailhol2-1/+14
The "--empty" option accepts "drop", "keep", or "abort" for the "drop" and "fixup" subcommands. Complete these values for the documented --empty=<value> form. While parse-options also accepts the split --empty <value> form, it is not documented. Omit it from completion as a trade-off for code simplicity. Signed-off-by: Vincent Mailhol <mailhol@kernel.org> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13completion: add 'git history' subcommandsVincent Mailhol2-0/+75
Use the parse-options completion helpers for the git history subcommands and their options. All current history subcommands take a revision as their first positional argument, so complete that argument as a revision. Once the revision is present, leave any further positional arguments to subcommand-specific completion. This allows a subcommand to complete another kind of argument, such as the pathspec accepted by git history split or another revision if a future subcommand accepts one. Signed-off-by: Vincent Mailhol <mailhol@kernel.org> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13completion: 'git checkout' completes untracked paths as a last resortJunio C Hamano2-2/+23
We taught 'git checkout' to first try to complete revisions (unless '--' is present on the command line) and, failing that, to complete tracked paths. If this yields nothing, it lets the Bash default, which offers paths in $PWD, kick in. Teach it to complete untracked paths before giving up and letting the Bash default kick in. With this change, $ git -C another-directory checkout un<TAB> finds the 'untracked' file in another-directory and offers it as a completion candidate. Note that this is of somewhat dubious value, as an untracked path by definition does not exist in the index, so checking it out from the index would not work well. Even when used to check out the path from a different branch, it is still of dubious value because it is unlikely that a path tracked in another branch is lying untracked in the working tree, as switching from a branch with the path to a branch without it will normally remove the file in the working tree. A better behavior probably is to detect the tree-ish argument on the command line and offer paths with the given prefix as candidates, but there is no __git_complete_from_tree() helper readily usable, so mark this as #leftoverbits to wait for another day. Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13completion: complete tracked paths for "git checkout"Junio C Hamano2-0/+43
When completing arguments for "git checkout", _git_checkout() delegates to __git_complete_refs(), which only completes revision references. This is good, as mixing revisions and paths in a single list from which the user can choose is confusing. However, if no reference matches, or if "--" is given, _git_checkout() leaves COMPREPLY empty. Bash then falls back to the default filename completion in $PWD. This fails when "git -C <path>" is used, as $PWD is not the target repository. Update _git_checkout() to use __git_complete_index_file() when "--" is present, or when revision reference completion yields no matching candidates, so that tracked paths are offered as candidates. Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13completion: no-op refactoring of checkout completionJunio C Hamano1-40/+42
The 'git checkout' completion function punts very early when it sees '--' on the command line, as it indicates that options or revisions can no longer appear. By returning early, it allows the default Bash action (which completes files in '$PWD') to kick in. In preparation for changing what happens in the next step when option or revision completion yields no matching candidates, or when '--' is present, reorganize the control flow to avoid this early return, and add explicit returns to the option completion branches. Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13packfile: fix perf regression with many packsJohannes Schindelin4-4/+10
Since 589127caa730 (packfile: move list of packs into the packfile store, 2025-10-30), there is a performance regression when many packfiles need to be loaded: `packfile_store_add_pack()` now calls `packfile_list_remove_internal()` to detect whether the packfile was _already_ in the list, and if so, move it to the end of the list. This function linearly scans the existing list before every insertion. Newly loading N packs therefore has complexity O(N²). In one reported use case (https://github.com/microsoft/git/issues/970), N equals 37,815 and caused a slow-down of a simple `git rev-parse --short HEAD` (which is regularly executed as part of `GIT_PS1`) from 0.4s to 4.5s. Let's fix this by establishing a fast path for known-new packfiles. The keen reader will note that there is currently only a single, "known-new" caller of the `packfile_list_append()` function, and wonder why not simply remove this check whether the packfile already exists in the list? Originally, when above-mentioned commit introduced that logic, there was a second caller in `prepare_midx()`, which would have required that check, but that caller was removed in 6aff1f25a046 (packfile: always add packfiles to MRU when adding a pack, 2025-10-30). Still, the function is declared in a header file, and to avoid any problems with in-flight or downstream callers, it is safer to extend the signature to be explicit whether or not to skip that check. Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13packfile: widen `unpack_object_header_buffer()` to `size_t`Johannes Schindelin4-12/+9
As part of the ongoing effort to replace `unsigned long` data types with `size_t` wherever appropriate (mainly to fix all those problems on Windows with objects larger than 4GB), let's also adjust the return type and the type of the `len` parameter of this function. Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13git-zlib: widen `git_deflate_bound()` to `size_t`Johannes Schindelin2-3/+15
All four `unsigned long`/`int`/`ssize_t` receivers across archive-zip, diff, http-push and t/helper/test-pack-deltas were widened to `size_t` in the prior commits, and remote-curl and fast-import were already there. With every caller prepared, both the parameter and the return type can now move without introducing any silent narrowing. For inputs above zlib's `uLong` range (i.e. >4 GiB on platforms where `uLong` is 32-bit, notably 64-bit Windows), defer to zlib's stored-block formula (the same fallback it would itself use, see https://github.com/madler/zlib/blob/v1.3.2/deflate.c#L832-L928 keeping in mind that for large sizes, the `storelen` would be relevant, also compare with https://github.com/madler/zlib/issues/549 for a fuller story) plus the worst-case wrapper overhead. The existing path through `deflateBound()` is unchanged for inputs that fit. Assisted-by: Opus 4.7 Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13t/helper/test-pack-deltas: widen `do_compress()`'s maxsize local to `size_t`Johannes Schindelin1-1/+1
Prep for the upcoming `git_deflate_bound()` widening to `size_t`. The local is only ever the return value of `git_deflate_bound()` and the `xmalloc()`/`stream.avail_out` sizes derived from it; widening it has no semantic effect today. Assisted-by: Opus 4.7 Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de> Signed-off-by: Junio C Hamano <gitster@pobox.com>
2026-08-13http-push: widen `start_put()`'s size local from `ssize_t` to `size_t`Johannes Schindelin1-1/+1
The local is initialised from `git_deflate_bound()` (an unsigned upper bound on the deflated output, never negative) and used in exactly three places: the initialising assignment, `strbuf_grow(buf, size)` whose parameter is already `size_t`, and `stream.avail_out` which became `size_t` in the prior commit. There is no comparison against zero or a negative value, no subtraction, no arithmetic that depends on signedness, and no path that would assign a signed quantity to it. The original `ssize_t` was the wrong type to begin with: a `git_deflate_bound()` result above `SSIZE_MAX` would have wrapped negative on assignment and then implicitly re-extended to a huge `size_t` at `strbuf_grow()`/`stream.avail_out`, requesting an absurd allocation. That is not a real-world concern for the object sizes http-push pushes today, but it is also the reason the type needs to move to `size_t` before `git_deflate_bound()` itself is widened. Assisted-by: Opus 4.7 Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de> Signed-off-by: Junio C Hamano <gitster@pobox.com>