diff --git a/research/code-review/img/pr_bot_vs_human_comparison.png b/research/code-review/img/pr_bot_vs_human_comparison.png new file mode 100644 index 0000000..40fa058 Binary files /dev/null and b/research/code-review/img/pr_bot_vs_human_comparison.png differ diff --git a/research/code-review/img/pr_distribution_buckets.png b/research/code-review/img/pr_distribution_buckets.png new file mode 100644 index 0000000..7c2a759 Binary files /dev/null and b/research/code-review/img/pr_distribution_buckets.png differ diff --git a/research/code-review/img/pr_time_distributions.png b/research/code-review/img/pr_time_distributions.png new file mode 100644 index 0000000..8d5aff6 Binary files /dev/null and b/research/code-review/img/pr_time_distributions.png differ diff --git a/research/code-review/img/pr_time_distributions_zoomed.png b/research/code-review/img/pr_time_distributions_zoomed.png new file mode 100644 index 0000000..789cc91 Binary files /dev/null and b/research/code-review/img/pr_time_distributions_zoomed.png differ diff --git a/research/code-review/index.md b/research/code-review/index.md new file mode 100644 index 0000000..4dfd512 --- /dev/null +++ b/research/code-review/index.md @@ -0,0 +1,94 @@ +--- +title: Code review improvements in Packit (+ automation) +authors: mfocko +--- + +This analysis looks at pull requests in the Packit organization over a rolling 365-day period. It focuses on three timestamps: + +- when a pull request was opened; +- when the first review was submitted, separating human and bot reviewers; and +- when the pull request was merged. + +The source data contains 1,138 merged pull requests. Of these, 1,126 had a recorded review, 1,118 had a human review, and 813 had a bot review. The scripts used to calculate the timings are included alongside this post. + +## The headline result + +Most review activity happens quickly, but the averages are pulled upward by a small number of very old pull requests. The median time to the first human review is about 30 minutes, while the median time to merge is about 15 hours. The corresponding means are much larger: 1.68 days for a human review and 5.17 days to merge. + +That difference is not a contradiction. It says that a typical pull request moves quickly, while a minority remains open for days or weeks. For this workflow, medians and distributions are more useful operational measures than means alone. + +![Distributions of merge and review timings](img/pr_time_distributions.png) + +## First review: automation is almost immediate + +The first-review distribution is concentrated at the left edge of the chart: + +- 974 of 1,126 pull requests, or 86.5%, received some kind of review within one hour. +- 39 (3.5%) received their first review in one to six hours. +- 38 (3.4%) took six to 24 hours. +- Only 75 pull requests, 6.7%, took more than a day to receive a first review. + +The overall median is effectively zero days and the mean is 1.04 days. This is another example of a long tail: the majority are reviewed immediately, but a few late reviews dominate the average. + +The bot and human measurements explain why the first-review number is so low. Bot reviews have a median of roughly zero hours and a mean of 0.17 days, or about four hours. Human reviews have a median of about 0.5 hours and a mean of 1.68 days. + +![Bot and human review timings compared](img/pr_bot_vs_human_comparison.png) + +The bucket view makes the difference especially clear: + +- 805 of 813 pull requests with a bot review, 99.0%, received it within one hour. +- 684 of 1,118 pull requests with a human review, 61.2%, received it within one hour. +- A further 275 human reviews, 24.6%, arrived between one and 24 hours. +- Human reviews still had a visible tail: 104, or 9.3%, arrived after three days. + +The automation is therefore doing something valuable even when it does not replace a human reviewer: it provides immediate feedback while the pull request is waiting for a person. + +![Review timing buckets](img/pr_distribution_buckets.png) + +## Merge time is a different metric + +Review arrival and merge completion should not be treated as the same outcome. A review can be available immediately while a pull request waits for changes, discussion, a maintainer decision, CI, or an appropriate merge window. + +The merge-time distribution has a median of 0.62 days, approximately 15 hours, and a mean of 5.17 days. Its buckets are: + +- 278 pull requests (24.4%) merged in less than one hour. +- 230 (20.2%) merged in one to six hours. +- 214 (18.8%) merged in six to 24 hours. +- 282 (24.8%) merged between one and seven days. +- 134 (11.8%) took more than one week. +- 43 (3.8%) took more than four weeks. + +In other words, 63.4% merged within a day, but more than one in ten remained open for over a week. The long tail is large enough to matter to planning even though it does not describe the typical pull request. + +![The first 24 hours in more detail](img/pr_time_distributions_zoomed.png) + +The zoomed view also shows that fast merges are not all clustered at one instant. There is a substantial set of merges in the first hour, followed by a gradual decline throughout the day. This suggests that both immediate automation and human availability contribute to the observed throughput. + +## What this suggests for our process + +### Keep automation early and non-blocking where appropriate + +Bot feedback arrives faster and more consistently than human feedback. That makes it well suited to cheap checks, dependency updates, formatting, and other feedback that benefits from being available before a human starts reviewing. The data does not show that bot feedback is sufficient for every change, but it does show that it can shorten the time to useful first feedback. + +### Optimize the tail, not just the median + +The median experience is already fairly fast. The larger opportunity is the tail: pull requests that receive no timely human attention or that remain open for more than a week. A useful next step would be to identify the causes of those cases, such as ownership gaps, review requests without a response, failing checks, author inactivity, or changes that require cross-team coordination. + +### Track review latency separately from delivery latency + +"Time to first review" measures responsiveness. "Time to merge" measures the whole path to integration. Combining them into one metric would hide where delay occurs. Both should be reported, preferably with medians and percentile or bucket distributions rather than only means. + +## Method and limitations + +The analysis scripts read one JSON file per pull request, calculate elapsed time from `createdAt` to review and merge timestamps, classify reviewers as bots or humans, and generate the figures in this directory. The bot classification uses a known-account list plus login-name heuristics such as `-bot`, `[bot]`, or the substring `bot`. + +There are several important limitations: + +- The sample is limited to merged pull requests, so abandoned or still-open work is not represented. +- The figures report available observations for each metric, not necessarily the same set of pull requests. For example, the plotted samples are 1,138 for merge time, 1,126 for any review, 1,118 for human review, and 813 for bot review. +- The scripts use the first review in the API response for the basic analysis. This assumes the review list is ordered chronologically; the bot-aware analysis explicitly scans timestamps when separating humans and bots. +- A review submission is only a proxy for useful feedback. It does not measure review depth, requested changes, discussion quality, or whether the author acted on the review. +- The percentile implementation uses indexed observations rather than an interpolated percentile definition, so percentile values should be treated as approximate. +- The one-year snapshot describes this period and organization. It is not a universal benchmark for other projects. + +These caveats do not change the main conclusion: Packit's review process is fast for the typical pull request, automation provides feedback exceptionally quickly, and the main source of delay is the smaller population in the long tail. Future analyses should connect timing with pull-request size, repository, author and reviewer workload, CI status, and outcome so that improvements target causes rather than just symptoms. diff --git a/research/code-review/scripts/analyze_pr_timings.jq b/research/code-review/scripts/analyze_pr_timings.jq new file mode 100644 index 0000000..2d39de8 --- /dev/null +++ b/research/code-review/scripts/analyze_pr_timings.jq @@ -0,0 +1,29 @@ +# Calculate time differences in hours +def time_diff_hours(from; to): + if from and to then + ((to | fromdateiso8601) - (from | fromdateiso8601)) / 3600 + else + null + end; + +# Calculate time differences in days +def time_diff_days(from; to): + if from and to then + ((to | fromdateiso8601) - (from | fromdateiso8601)) / 86400 + else + null + end; + +# Process each PR +{ + repo, + number, + created: .createdAt, + merged: .mergedAt, + first_review: (.reviews[0].submittedAt // null), + review_count: (.reviews | length), + time_to_first_review_hours: time_diff_hours(.createdAt; .reviews[0].submittedAt // null), + time_to_first_review_days: time_diff_days(.createdAt; .reviews[0].submittedAt // null), + time_to_merge_hours: time_diff_hours(.createdAt; .mergedAt), + time_to_merge_days: time_diff_days(.createdAt; .mergedAt) +} diff --git a/research/code-review/scripts/analyze_timings.py b/research/code-review/scripts/analyze_timings.py new file mode 100755 index 0000000..708e3ef --- /dev/null +++ b/research/code-review/scripts/analyze_timings.py @@ -0,0 +1,215 @@ +#!/usr/bin/env python3 +import json +import os +from datetime import datetime +from pathlib import Path +import statistics + + +def parse_timestamp(ts): + """Parse ISO timestamp to datetime""" + if not ts: + return None + return datetime.fromisoformat(ts.replace("Z", "+00:00")) + + +def calculate_hours(start, end): + """Calculate hours between two timestamps""" + if not start or not end: + return None + return (end - start).total_seconds() / 3600 + + +def calculate_days(start, end): + """Calculate days between two timestamps""" + if not start or not end: + return None + return (end - start).total_seconds() / 86400 + + +def process_pr_file(filepath): + """Process a single PR JSON file""" + try: + with open(filepath) as f: + data = json.load(f) + + # Extract repo and number from filename + filename = Path(filepath).stem + parts = filename.rsplit("_", 1) + repo = parts[0].replace("_", "/", 1) + number = int(parts[1]) + + created = parse_timestamp(data.get("createdAt")) + merged = parse_timestamp(data.get("mergedAt")) + reviews = data.get("reviews", []) + + first_review = None + if reviews: + first_review = parse_timestamp(reviews[0].get("submittedAt")) + + result = { + "repo": repo, + "number": number, + "created": data.get("createdAt"), + "merged": data.get("mergedAt"), + "review_count": len(reviews), + "has_review": len(reviews) > 0, + "time_to_first_review_hours": calculate_hours(created, first_review), + "time_to_first_review_days": calculate_days(created, first_review), + "time_to_merge_hours": calculate_hours(created, merged), + "time_to_merge_days": calculate_days(created, merged), + } + + return result + except Exception as e: + print(f"Error processing {filepath}: {e}") + return None + + +def main(): + data_dir = "/tmp/pr_data" + results = [] + + print("Processing PR data files...") + for filepath in sorted(Path(data_dir).glob("*.json")): + result = process_pr_file(filepath) + if result: + results.append(result) + + print(f"\nProcessed {len(results)} PRs") + + # Save results + with open("/tmp/pr_timing_analysis.json", "w") as f: + json.dump(results, f, indent=2) + + # Calculate statistics + time_to_merge = [ + r["time_to_merge_days"] for r in results if r["time_to_merge_days"] is not None + ] + time_to_review = [ + r["time_to_first_review_days"] + for r in results + if r["time_to_first_review_days"] is not None + ] + + print("\n=== TIME TO MERGE STATISTICS ===") + if time_to_merge: + print(f"Total PRs with merge data: {len(time_to_merge)}") + print(f"Mean: {statistics.mean(time_to_merge):.2f} days") + print(f"Median: {statistics.median(time_to_merge):.2f} days") + print(f"Min: {min(time_to_merge):.2f} days") + print(f"Max: {max(time_to_merge):.2f} days") + if len(time_to_merge) > 1: + print(f"Std Dev: {statistics.stdev(time_to_merge):.2f} days") + + # Percentiles + sorted_merge = sorted(time_to_merge) + p25 = sorted_merge[len(sorted_merge) // 4] + p75 = sorted_merge[3 * len(sorted_merge) // 4] + p90 = sorted_merge[9 * len(sorted_merge) // 10] + p95 = sorted_merge[95 * len(sorted_merge) // 100] + + print(f"25th percentile: {p25:.2f} days") + print(f"75th percentile: {p75:.2f} days") + print(f"90th percentile: {p90:.2f} days") + print(f"95th percentile: {p95:.2f} days") + + print("\n=== TIME TO FIRST REVIEW STATISTICS ===") + if time_to_review: + print(f"Total PRs with review data: {len(time_to_review)}") + print(f"Mean: {statistics.mean(time_to_review):.2f} days") + print(f"Median: {statistics.median(time_to_review):.2f} days") + print(f"Min: {min(time_to_review):.2f} days") + print(f"Max: {max(time_to_review):.2f} days") + if len(time_to_review) > 1: + print(f"Std Dev: {statistics.stdev(time_to_review):.2f} days") + + # Percentiles + sorted_review = sorted(time_to_review) + p25 = sorted_review[len(sorted_review) // 4] + p75 = sorted_review[3 * len(sorted_review) // 4] + p90 = sorted_review[9 * len(sorted_review) // 10] + p95 = sorted_review[95 * len(sorted_review) // 100] + + print(f"25th percentile: {p25:.2f} days") + print(f"75th percentile: {p75:.2f} days") + print(f"90th percentile: {p90:.2f} days") + print(f"95th percentile: {p95:.2f} days") + + prs_without_review = sum(1 for r in results if not r["has_review"]) + print(f"\n=== REVIEW STATUS ===") + print(f"PRs with reviews: {len(time_to_review)}/{len(results)}") + print(f"PRs without reviews: {prs_without_review}/{len(results)}") + + # Distribution buckets for merge time + print("\n=== TIME TO MERGE DISTRIBUTION ===") + buckets = { + "< 1 hour": 0, + "1-6 hours": 0, + "6-24 hours": 0, + "1-3 days": 0, + "3-7 days": 0, + "1-2 weeks": 0, + "2-4 weeks": 0, + "> 4 weeks": 0, + } + + for days in time_to_merge: + hours = days * 24 + if hours < 1: + buckets["< 1 hour"] += 1 + elif hours < 6: + buckets["1-6 hours"] += 1 + elif hours < 24: + buckets["6-24 hours"] += 1 + elif days < 3: + buckets["1-3 days"] += 1 + elif days < 7: + buckets["3-7 days"] += 1 + elif days < 14: + buckets["1-2 weeks"] += 1 + elif days < 28: + buckets["2-4 weeks"] += 1 + else: + buckets["> 4 weeks"] += 1 + + for bucket, count in buckets.items(): + pct = (count / len(time_to_merge) * 100) if time_to_merge else 0 + print(f"{bucket:15s}: {count:4d} ({pct:5.1f}%)") + + # Distribution buckets for review time + print("\n=== TIME TO FIRST REVIEW DISTRIBUTION ===") + review_buckets = { + "< 1 hour": 0, + "1-6 hours": 0, + "6-24 hours": 0, + "1-3 days": 0, + "3-7 days": 0, + "1-2 weeks": 0, + "> 2 weeks": 0, + } + + for days in time_to_review: + hours = days * 24 + if hours < 1: + review_buckets["< 1 hour"] += 1 + elif hours < 6: + review_buckets["1-6 hours"] += 1 + elif hours < 24: + review_buckets["6-24 hours"] += 1 + elif days < 3: + review_buckets["1-3 days"] += 1 + elif days < 7: + review_buckets["3-7 days"] += 1 + elif days < 14: + review_buckets["1-2 weeks"] += 1 + else: + review_buckets["> 2 weeks"] += 1 + + for bucket, count in review_buckets.items(): + pct = (count / len(time_to_review) * 100) if time_to_review else 0 + print(f"{bucket:15s}: {count:4d} ({pct:5.1f}%)") + + +if __name__ == "__main__": + main() diff --git a/research/code-review/scripts/analyze_timings_with_bots.py b/research/code-review/scripts/analyze_timings_with_bots.py new file mode 100755 index 0000000..72cd3b4 --- /dev/null +++ b/research/code-review/scripts/analyze_timings_with_bots.py @@ -0,0 +1,309 @@ +#!/usr/bin/env python3 +import json +import os +from datetime import datetime +from pathlib import Path +import statistics + +# Known bot accounts +BOT_ACCOUNTS = { + "gemini-code-assist", + "fullsend-ai-coder", + "dependabot", + "pre-commit-ci", + "github-actions", +} + + +def is_bot_reviewer(login): + """Check if a reviewer is a bot""" + if not login: + return False + login_lower = login.lower() + # Check if it's a known bot or ends with [bot] or -bot + return ( + login in BOT_ACCOUNTS + or login_lower.endswith("[bot]") + or login_lower.endswith("-bot") + or "bot" in login_lower + ) + + +def parse_timestamp(ts): + """Parse ISO timestamp to datetime""" + if not ts: + return None + return datetime.fromisoformat(ts.replace("Z", "+00:00")) + + +def calculate_hours(start, end): + """Calculate hours between two timestamps""" + if not start or not end: + return None + return (end - start).total_seconds() / 3600 + + +def calculate_days(start, end): + """Calculate days between two timestamps""" + if not start or not end: + return None + return (end - start).total_seconds() / 86400 + + +def process_pr_file(filepath): + """Process a single PR JSON file""" + try: + with open(filepath) as f: + data = json.load(f) + + # Extract repo and number from filename + filename = Path(filepath).stem + parts = filename.rsplit("_", 1) + repo = parts[0].replace("_", "/", 1) + number = int(parts[1]) + + created = parse_timestamp(data.get("createdAt")) + merged = parse_timestamp(data.get("mergedAt")) + reviews = data.get("reviews", []) + + first_review = None + first_review_by = None + first_human_review = None + first_bot_review = None + human_review_count = 0 + bot_review_count = 0 + + for review in reviews: + reviewer_login = review.get("author", {}).get("login", "") + review_time = parse_timestamp(review.get("submittedAt")) + + if not review_time: + continue + + # Track first review overall + if first_review is None: + first_review = review_time + first_review_by = reviewer_login + + # Track bot vs human reviews + if is_bot_reviewer(reviewer_login): + bot_review_count += 1 + if first_bot_review is None: + first_bot_review = review_time + else: + human_review_count += 1 + if first_human_review is None: + first_human_review = review_time + + result = { + "repo": repo, + "number": number, + "created": data.get("createdAt"), + "merged": data.get("mergedAt"), + "total_review_count": len(reviews), + "human_review_count": human_review_count, + "bot_review_count": bot_review_count, + "has_review": len(reviews) > 0, + "has_human_review": human_review_count > 0, + "has_bot_review": bot_review_count > 0, + "first_review_by": first_review_by, + "first_review_is_bot": ( + is_bot_reviewer(first_review_by) if first_review_by else None + ), + "time_to_first_review_hours": calculate_hours(created, first_review), + "time_to_first_review_days": calculate_days(created, first_review), + "time_to_first_human_review_hours": calculate_hours( + created, first_human_review + ), + "time_to_first_human_review_days": calculate_days( + created, first_human_review + ), + "time_to_first_bot_review_hours": calculate_hours( + created, first_bot_review + ), + "time_to_first_bot_review_days": calculate_days(created, first_bot_review), + "time_to_merge_hours": calculate_hours(created, merged), + "time_to_merge_days": calculate_days(created, merged), + } + + return result + except Exception as e: + print(f"Error processing {filepath}: {e}") + return None + + +def print_distribution(title, time_data, buckets_def): + """Print distribution buckets""" + print(f"\n=== {title} ===") + buckets = {k: 0 for k in buckets_def.keys()} + + for days in time_data: + hours = days * 24 + for bucket_name, (min_hours, max_hours) in buckets_def.items(): + if min_hours <= hours < max_hours: + buckets[bucket_name] += 1 + break + + for bucket, count in buckets.items(): + pct = (count / len(time_data) * 100) if time_data else 0 + print(f"{bucket:15s}: {count:4d} ({pct:5.1f}%)") + + +def main(): + data_dir = "/tmp/pr_data" + results = [] + + print("Processing PR data files...") + for filepath in sorted(Path(data_dir).glob("*.json")): + result = process_pr_file(filepath) + if result: + results.append(result) + + print(f"\nProcessed {len(results)} PRs") + + # Save results + with open("/tmp/pr_timing_analysis_detailed.json", "w") as f: + json.dump(results, f, indent=2) + + # Separate data by review type + time_to_merge = [ + r["time_to_merge_days"] for r in results if r["time_to_merge_days"] is not None + ] + time_to_any_review = [ + r["time_to_first_review_days"] + for r in results + if r["time_to_first_review_days"] is not None + ] + time_to_human_review = [ + r["time_to_first_human_review_days"] + for r in results + if r["time_to_first_human_review_days"] is not None + ] + time_to_bot_review = [ + r["time_to_first_bot_review_days"] + for r in results + if r["time_to_first_bot_review_days"] is not None + ] + + # Calculate statistics + def print_stats(title, data): + print(f"\n{'='*60}") + print(f"{title}") + print("=" * 60) + if data: + print(f"Total PRs: {len(data)}") + print( + f"Mean: {statistics.mean(data):.2f} days ({statistics.mean(data)*24:.1f} hours)" + ) + print( + f"Median: {statistics.median(data):.2f} days ({statistics.median(data)*24:.1f} hours)" + ) + print(f"Min: {min(data):.2f} days ({min(data)*24:.1f} hours)") + print(f"Max: {max(data):.2f} days ({max(data)*24:.1f} hours)") + if len(data) > 1: + print(f"Std Dev: {statistics.stdev(data):.2f} days") + + # Percentiles + sorted_data = sorted(data) + p25 = sorted_data[len(sorted_data) // 4] + p75 = sorted_data[3 * len(sorted_data) // 4] + p90 = sorted_data[9 * len(sorted_data) // 10] + p95 = sorted_data[95 * len(sorted_data) // 100] + + print(f"\nPercentiles:") + print(f" 25th: {p25:.2f} days ({p25*24:.1f} hours)") + print(f" 75th: {p75:.2f} days ({p75*24:.1f} hours)") + print(f" 90th: {p90:.2f} days ({p90*24:.1f} hours)") + print(f" 95th: {p95:.2f} days ({p95*24:.1f} hours)") + else: + print("No data available") + + print_stats("TIME TO MERGE", time_to_merge) + print_stats("TIME TO FIRST REVIEW (ANY)", time_to_any_review) + print_stats("TIME TO FIRST HUMAN REVIEW", time_to_human_review) + print_stats("TIME TO FIRST BOT REVIEW", time_to_bot_review) + + # Review status breakdown + prs_with_human_review = sum(1 for r in results if r["has_human_review"]) + prs_with_bot_review = sum(1 for r in results if r["has_bot_review"]) + prs_without_review = sum(1 for r in results if not r["has_review"]) + prs_first_review_bot = sum(1 for r in results if r["first_review_is_bot"]) + prs_first_review_human = sum( + 1 for r in results if r["first_review_is_bot"] == False + ) + + print(f"\n{'='*60}") + print("REVIEW STATUS BREAKDOWN") + print("=" * 60) + print(f"Total PRs: {len(results)}") + print( + f"PRs with human reviews: {prs_with_human_review} ({prs_with_human_review/len(results)*100:.1f}%)" + ) + print( + f"PRs with bot reviews: {prs_with_bot_review} ({prs_with_bot_review/len(results)*100:.1f}%)" + ) + print( + f"PRs without any review: {prs_without_review} ({prs_without_review/len(results)*100:.1f}%)" + ) + print( + f"\nFirst review by bot: {prs_first_review_bot} ({prs_first_review_bot/len(results)*100:.1f}%)" + ) + print( + f"First review by human: {prs_first_review_human} ({prs_first_review_human/len(results)*100:.1f}%)" + ) + + # Distribution buckets + merge_buckets = { + "< 1 hour": (0, 1), + "1-6 hours": (1, 6), + "6-24 hours": (6, 24), + "1-3 days": (24, 72), + "3-7 days": (72, 168), + "1-2 weeks": (168, 336), + "2-4 weeks": (336, 672), + "> 4 weeks": (672, float("inf")), + } + + review_buckets = { + "< 1 hour": (0, 1), + "1-6 hours": (1, 6), + "6-24 hours": (6, 24), + "1-3 days": (24, 72), + "3-7 days": (72, 168), + "1-2 weeks": (168, 336), + "> 2 weeks": (336, float("inf")), + } + + print_distribution("TIME TO MERGE DISTRIBUTION", time_to_merge, merge_buckets) + print_distribution( + "TIME TO FIRST REVIEW (ANY) DISTRIBUTION", time_to_any_review, review_buckets + ) + print_distribution( + "TIME TO FIRST HUMAN REVIEW DISTRIBUTION", time_to_human_review, review_buckets + ) + print_distribution( + "TIME TO FIRST BOT REVIEW DISTRIBUTION", time_to_bot_review, review_buckets + ) + + # Top bot reviewers + print(f"\n{'='*60}") + print("BOT REVIEWERS") + print("=" * 60) + bot_reviewers = {} + for filepath in Path(data_dir).glob("*.json"): + try: + with open(filepath) as f: + data = json.load(f) + for review in data.get("reviews", []): + reviewer = review.get("author", {}).get("login", "") + if is_bot_reviewer(reviewer): + bot_reviewers[reviewer] = bot_reviewers.get(reviewer, 0) + 1 + except: + pass + + for bot, count in sorted(bot_reviewers.items(), key=lambda x: x[1], reverse=True): + print(f"{bot:30s}: {count:4d} reviews") + + +if __name__ == "__main__": + main() diff --git a/research/code-review/scripts/create_histograms.py b/research/code-review/scripts/create_histograms.py new file mode 100755 index 0000000..e58582d --- /dev/null +++ b/research/code-review/scripts/create_histograms.py @@ -0,0 +1,408 @@ +#!/usr/bin/env python3 +import json +import matplotlib.pyplot as plt +import matplotlib.patches as mpatches +import numpy as np +from pathlib import Path + +# Load the analysis data +with open("/tmp/pr_timing_analysis_detailed.json") as f: + data = json.load(f) + +# Extract time data +time_to_merge = [ + r["time_to_merge_days"] for r in data if r["time_to_merge_days"] is not None +] +time_to_any_review = [ + r["time_to_first_review_days"] + for r in data + if r["time_to_first_review_days"] is not None +] +time_to_human_review = [ + r["time_to_first_human_review_days"] + for r in data + if r["time_to_first_human_review_days"] is not None +] +time_to_bot_review = [ + r["time_to_first_bot_review_days"] + for r in data + if r["time_to_first_bot_review_days"] is not None +] + +# Create figure with 2x2 subplots +fig, ((ax1, ax2), (ax3, ax4)) = plt.subplots(2, 2, figsize=(16, 12)) +fig.suptitle( + "PR Time Distributions - Packit Organization (Last 365 Days)", + fontsize=16, + fontweight="bold", +) + +# Color scheme +color_merge = "#3498db" +color_any = "#9b59b6" +color_human = "#e74c3c" +color_bot = "#2ecc71" + + +def create_histogram(ax, data, title, color, xlabel, bins=50, xlim=None): + """Create a histogram with statistics""" + n, bins_edges, patches = ax.hist( + data, bins=bins, color=color, alpha=0.7, edgecolor="black", linewidth=0.5 + ) + + # Add median and mean lines + median = np.median(data) + mean = np.mean(data) + + ax.axvline( + median, color="red", linestyle="--", linewidth=2, label=f"Median: {median:.2f}d" + ) + ax.axvline( + mean, color="orange", linestyle="--", linewidth=2, label=f"Mean: {mean:.2f}d" + ) + + ax.set_xlabel(xlabel, fontsize=11, fontweight="bold") + ax.set_ylabel("Number of PRs", fontsize=11, fontweight="bold") + ax.set_title(title, fontsize=13, fontweight="bold", pad=10) + ax.legend(loc="upper right", fontsize=10) + ax.grid(True, alpha=0.3, linestyle="--") + + if xlim: + ax.set_xlim(xlim) + + # Add statistics text box + stats_text = f"n={len(data)}\nMin: {min(data):.2f}d\nMax: {max(data):.2f}d\nStd: {np.std(data):.2f}d" + ax.text( + 0.98, + 0.97, + stats_text, + transform=ax.transAxes, + fontsize=9, + verticalalignment="top", + horizontalalignment="right", + bbox=dict(boxstyle="round", facecolor="wheat", alpha=0.5), + ) + + return ax + + +# 1. Time to Merge +create_histogram( + ax1, + time_to_merge, + f"Time to Merge (n={len(time_to_merge)})", + color_merge, + "Days from PR Creation to Merge", + bins=60, + xlim=(0, min(60, max(time_to_merge))), +) + +# 2. Time to First Review (Any) +create_histogram( + ax2, + time_to_any_review, + f"Time to First Review - Any (n={len(time_to_any_review)})", + color_any, + "Days from PR Creation to First Review", + bins=50, + xlim=(0, min(30, max(time_to_any_review))), +) + +# 3. Time to First Human Review +create_histogram( + ax3, + time_to_human_review, + f"Time to First Human Review (n={len(time_to_human_review)})", + color_human, + "Days from PR Creation to First Human Review", + bins=50, + xlim=(0, min(30, max(time_to_human_review))), +) + +# 4. Time to First Bot Review +create_histogram( + ax4, + time_to_bot_review, + f"Time to First Bot Review (n={len(time_to_bot_review)})", + color_bot, + "Days from PR Creation to First Bot Review", + bins=50, + xlim=(0, min(10, max(time_to_bot_review))), +) + +plt.tight_layout() +plt.savefig("/tmp/pr_time_distributions.png", dpi=300, bbox_inches="tight") +print("Saved: /tmp/pr_time_distributions.png") + +# Create a second figure with zoomed-in views (first 24 hours) +fig2, ((ax5, ax6), (ax7, ax8)) = plt.subplots(2, 2, figsize=(16, 12)) +fig2.suptitle( + "PR Time Distributions - First 24 Hours (Zoomed In)", fontsize=16, fontweight="bold" +) + + +def create_zoomed_histogram(ax, data, title, color, xlabel, max_hours=24): + """Create a histogram showing only first 24 hours in hours""" + # Convert to hours and filter + data_hours = [d * 24 for d in data if d * 24 <= max_hours] + + if not data_hours: + ax.text( + 0.5, + 0.5, + "No data in range", + ha="center", + va="center", + transform=ax.transAxes, + ) + return + + bins = np.arange(0, max_hours + 1, 1) # 1-hour bins + n, bins_edges, patches = ax.hist( + data_hours, bins=bins, color=color, alpha=0.7, edgecolor="black", linewidth=0.5 + ) + + # Add median and mean lines + if data_hours: + median = np.median(data_hours) + mean = np.mean(data_hours) + + ax.axvline( + median, + color="red", + linestyle="--", + linewidth=2, + label=f"Median: {median:.1f}h", + ) + ax.axvline( + mean, + color="orange", + linestyle="--", + linewidth=2, + label=f"Mean: {mean:.1f}h", + ) + + ax.set_xlabel(xlabel, fontsize=11, fontweight="bold") + ax.set_ylabel("Number of PRs", fontsize=11, fontweight="bold") + ax.set_title(title, fontsize=13, fontweight="bold", pad=10) + ax.legend(loc="upper right", fontsize=10) + ax.grid(True, alpha=0.3, linestyle="--") + ax.set_xlim(0, max_hours) + + # Add statistics text box + pct = len(data_hours) / len(data) * 100 + stats_text = f"{len(data_hours)}/{len(data)} PRs\n({pct:.1f}%)" + ax.text( + 0.98, + 0.97, + stats_text, + transform=ax.transAxes, + fontsize=10, + verticalalignment="top", + horizontalalignment="right", + bbox=dict(boxstyle="round", facecolor="wheat", alpha=0.5), + ) + + +create_zoomed_histogram( + ax5, + time_to_merge, + "Time to Merge (First 24 Hours)", + color_merge, + "Hours from PR Creation to Merge", +) + +create_zoomed_histogram( + ax6, + time_to_any_review, + "Time to First Review - Any (First 24 Hours)", + color_any, + "Hours from PR Creation to First Review", +) + +create_zoomed_histogram( + ax7, + time_to_human_review, + "Time to First Human Review (First 24 Hours)", + color_human, + "Hours from PR Creation to First Human Review", +) + +create_zoomed_histogram( + ax8, + time_to_bot_review, + "Time to First Bot Review (First 24 Hours)", + color_bot, + "Hours from PR Creation to First Bot Review", + max_hours=6, +) # Bot reviews are so fast, show only 6 hours + +plt.tight_layout() +plt.savefig("/tmp/pr_time_distributions_zoomed.png", dpi=300, bbox_inches="tight") +print("Saved: /tmp/pr_time_distributions_zoomed.png") + +# Create a third figure comparing bot vs human review times +fig3, ax9 = plt.subplots(1, 1, figsize=(14, 8)) + +# Create overlapping histograms +bins = np.arange(0, 10, 0.2) # 0-10 days in 0.2 day increments +ax9.hist( + time_to_bot_review, + bins=bins, + color=color_bot, + alpha=0.5, + label=f"Bot Reviews (n={len(time_to_bot_review)})", + edgecolor="black", + linewidth=0.5, +) +ax9.hist( + time_to_human_review, + bins=bins, + color=color_human, + alpha=0.5, + label=f"Human Reviews (n={len(time_to_human_review)})", + edgecolor="black", + linewidth=0.5, +) + +# Add median lines +bot_median = np.median(time_to_bot_review) +human_median = np.median(time_to_human_review) +ax9.axvline( + bot_median, + color=color_bot, + linestyle="--", + linewidth=2, + label=f"Bot Median: {bot_median:.3f}d ({bot_median*24:.1f}h)", +) +ax9.axvline( + human_median, + color=color_human, + linestyle="--", + linewidth=2, + label=f"Human Median: {human_median:.3f}d ({human_median*24:.1f}h)", +) + +ax9.set_xlabel("Days from PR Creation to First Review", fontsize=12, fontweight="bold") +ax9.set_ylabel("Number of PRs", fontsize=12, fontweight="bold") +ax9.set_title( + "Bot vs Human Review Time Comparison", fontsize=14, fontweight="bold", pad=15 +) +ax9.legend(loc="upper right", fontsize=11) +ax9.grid(True, alpha=0.3, linestyle="--") +ax9.set_xlim(0, 10) + +plt.tight_layout() +plt.savefig("/tmp/pr_bot_vs_human_comparison.png", dpi=300, bbox_inches="tight") +print("Saved: /tmp/pr_bot_vs_human_comparison.png") + +# Create stacked bar chart showing distribution buckets +fig4, ((ax10, ax11), (ax12, ax13)) = plt.subplots(2, 2, figsize=(16, 10)) +fig4.suptitle( + "Time Distribution Buckets - Percentage View", fontsize=16, fontweight="bold" +) + + +def create_bucket_chart(ax, data, title, color, buckets): + """Create a bar chart of time buckets""" + counts = {k: 0 for k in buckets.keys()} + + for days in data: + hours = days * 24 + for bucket_name, (min_h, max_h) in buckets.items(): + if min_h <= hours < max_h: + counts[bucket_name] += 1 + break + + labels = list(counts.keys()) + values = list(counts.values()) + percentages = [v / len(data) * 100 for v in values] + + bars = ax.bar( + range(len(labels)), + percentages, + color=color, + alpha=0.7, + edgecolor="black", + linewidth=1, + ) + ax.set_xticks(range(len(labels))) + ax.set_xticklabels(labels, rotation=45, ha="right") + ax.set_ylabel("Percentage of PRs (%)", fontsize=11, fontweight="bold") + ax.set_title(title, fontsize=12, fontweight="bold", pad=10) + ax.grid(True, alpha=0.3, linestyle="--", axis="y") + + # Add percentage labels on bars + for i, (bar, pct, count) in enumerate(zip(bars, percentages, values)): + if pct > 0: + ax.text( + bar.get_x() + bar.get_width() / 2, + bar.get_height() + 1, + f"{pct:.1f}%\n({count})", + ha="center", + va="bottom", + fontsize=8, + fontweight="bold", + ) + + +merge_buckets = { + "< 1h": (0, 1), + "1-6h": (1, 6), + "6-24h": (6, 24), + "1-3d": (24, 72), + "3-7d": (72, 168), + "1-2w": (168, 336), + "2-4w": (336, 672), + "> 4w": (672, float("inf")), +} + +review_buckets = { + "< 1h": (0, 1), + "1-6h": (1, 6), + "6-24h": (6, 24), + "1-3d": (24, 72), + "3-7d": (72, 168), + "1-2w": (168, 336), + "> 2w": (336, float("inf")), +} + +create_bucket_chart( + ax10, + time_to_merge, + f"Time to Merge (n={len(time_to_merge)})", + color_merge, + merge_buckets, +) +create_bucket_chart( + ax11, + time_to_any_review, + f"Time to Any Review (n={len(time_to_any_review)})", + color_any, + review_buckets, +) +create_bucket_chart( + ax12, + time_to_human_review, + f"Time to Human Review (n={len(time_to_human_review)})", + color_human, + review_buckets, +) +create_bucket_chart( + ax13, + time_to_bot_review, + f"Time to Bot Review (n={len(time_to_bot_review)})", + color_bot, + review_buckets, +) + +plt.tight_layout() +plt.savefig("/tmp/pr_distribution_buckets.png", dpi=300, bbox_inches="tight") +print("Saved: /tmp/pr_distribution_buckets.png") + +print("\nAll histograms created successfully!") +print("Files saved:") +print(" 1. /tmp/pr_time_distributions.png - Full distribution histograms") +print(" 2. /tmp/pr_time_distributions_zoomed.png - First 24 hours (zoomed)") +print(" 3. /tmp/pr_bot_vs_human_comparison.png - Bot vs Human overlay") +print(" 4. /tmp/pr_distribution_buckets.png - Bucket percentage charts")