Merge pull request #408 from Codium-ai/tr/final_fixes

fixed review
better link
2025-07-05 21:30:40 +08:00 · 2023-10-29 09:02:48 -07:00 · 2023-10-29 18:01:50 +02:00 · 2023-10-29 17:59:46 +02:00 · 2023-10-29 16:54:59 +02:00 · 2023-10-29 16:33:38 +02:00
20 changed files with 374 additions and 115 deletions
--- a/README.md
+++ b/README.md
@ -31,6 +31,8 @@ CodiumAI `PR-Agent` is an open-source tool aiming to help developers review pull
 ‣ **Find Similar Issue ([`/similar_issue`](./docs/SIMILAR_ISSUE.md))**: Automatically retrieves and presents similar issues
 \
 ‣ **Add Documentation ([`/add_docs`](./docs/ADD_DOCUMENTATION.md))**: Automatically adds documentation to un-documented functions/classes in the PR.
 \
 ‣ **Generate Custom Labels ([`/generate_labels`](./docs/GENERATE_CUSTOM_LABELS.md))**: Automatically suggests custom labels based on the PR code changes.
 See the [Installation Guide](./INSTALL.md) for instructions how to install and run the tool on different platforms.
@ -115,6 +117,7 @@ See the [Tools Guide](./docs/TOOLS_GUIDE.md) for detailed description of the dif
 |       | Update CHANGELOG.md                         |   :white_check_mark:    |   :white_check_mark:    |   :white_check_mark:        |   :white_check_mark:    |          |       |
 |       | Find similar issue                          |   :white_check_mark:    |                         |                             |          |          |       |
 |       | Add Documentation                           |   :white_check_mark:    |   :white_check_mark:    |   :white_check_mark:        |   :white_check_mark:    |          |    :white_check_mark:    |
 |       | Generate Labels                           |   :white_check_mark:    |   :white_check_mark:    |         |     |          |      |
 |       |                                             |        |        |      |      |      |
 | USAGE | CLI                                         |   :white_check_mark:    |   :white_check_mark:    |   :white_check_mark:       |   :white_check_mark:    |   :white_check_mark:    |
 |       | App / webhook                               |   :white_check_mark:    |   :white_check_mark:    |           |          |          |
--- a/RELEASE_NOTES.md
+++ b/RELEASE_NOTES.md
@ -1,3 +1,27 @@
 ## [Version 0.9] - 2023-10-29
 - codiumai/pr-agent:0.9
 - codiumai/pr-agent:0.9-github_app
 - codiumai/pr-agent:0.9-bitbucket-app
 - codiumai/pr-agent:0.9-gitlab_webhook
 - codiumai/pr-agent:0.9-github_polling
 - codiumai/pr-agent:0.9-github_action
 ### Added::Algo
 - New tool - [generate_labels](https://github.com/Codium-ai/pr-agent/blob/main/docs/GENERATE_CUSTOM_LABELS.md)
 - New ability to use [customize labels](https://github.com/Codium-ai/pr-agent/blob/main/docs/GENERATE_CUSTOM_LABELS.md#how-to-enable-custom-labels) on the `review` and `describe` tools.
 - GitHub Action: Can now use a `.pr_agent.toml` file to control configuration parameters (see [Usage Guide](./Usage.md#working-with-github-action)).
 - GitHub App: Added ability to trigger tools on [push events](https://github.com/Codium-ai/pr-agent/blob/main/Usage.md#github-app-automatic-tools-for-new-code-pr-push)
 - Support custom domain URLs for azure devops integration (see [link](https://github.com/Codium-ai/pr-agent/pull/381)).
 - PR Description default mode is now in [bullet points](https://github.com/Codium-ai/pr-agent/blob/main/pr_agent/settings/configuration.toml#L35).
 ### Added::Documentation
 Significant documentation updates (see [Installation Guide](https://github.com/Codium-ai/pr-agent/blob/main/INSTALL.md), [Usage Guide](https://github.com/Codium-ai/pr-agent/blob/main/Usage.md), and [Tools Guide](https://github.com/Codium-ai/pr-agent/blob/main/docs/TOOLS_GUIDE.md))
 ### Fixed
 - Fixed support for BitBucket pipeline (see [link](https://github.com/Codium-ai/pr-agent/pull/386))
 - Fixed a bug in `review -i` tool
 - Added blacklist for specific file extensions in `add_docs` tool (see [link](https://github.com/Codium-ai/pr-agent/pull/385/))
 ## [Version 0.8] - 2023-09-27
 - codiumai/pr-agent:0.8
 - codiumai/pr-agent:0.8-github_app
--- a/Usage.md
+++ b/Usage.md
@ -112,15 +112,27 @@ When running PR-Agent from [GitHub App](INSTALL.md#method-5-run-as-a-github-app)
 #### GitHub app automatic tools
 The [github_app](pr_agent/settings/configuration.toml#L56) section defines GitHub app specific configurations.  
-An important parameter is `pr_commands`, which is a list of tools that will be **run automatically** when a new PR is opened:
+In this section you can define configurations to control the conditions for which tools will **run automatically**.  
 Note that a local `.pr_agent.toml` file enables you to edit and customize the default parameters of any tool, not just the ones that are run automatically.
 ##### GitHub app automatic tools for PR actions
 The GitHub app can respond to the following actions on a PR:
 1. `opened` - Opening a new PR
 2. `reopened` - Reopening a closed PR
 3. `ready_for_review` - Moving a PR from Draft to Open
 4. `review_requested` - Specifically requesting review (in the PR reviewers list) from the `github-actions[bot]` user
 The configuration parameter `handle_pr_actions` defines the list of actions for which the GitHub app will trigger the PR-Agent.  
 The configuration parameter `pr_commands` defines the list of tools that will be **run automatically** when one of the above action happens (e.g. a new PR is opened):
 ```
 [github_app]
 handle_pr_actions = ['opened', 'reopened', 'ready_for_review', 'review_requested']
 pr_commands = [
    "/describe --pr_description.add_original_user_description=true --pr_description.keep_original_user_title=true",
    "/auto_review",
 ]
 ```
-This means that when a new PR is opened, PR-Agent will run the `describe` and `auto_review` tools.
+This means that when a new PR is opened/reopened or marked as ready for review, PR-Agent will run the `describe` and `auto_review` tools.  
 For the describe tool, the `add_original_user_description` and `keep_original_user_title` parameters will be set to true.
 You can override the default tool parameters by uploading a local configuration file called `.pr_agent.toml` to the root of your repo.
@ -135,11 +147,27 @@ When a new PR is opened, PR-Agent will run the `describe` tool with the above pa
 To cancel the automatic run of all the tools, set:
 ```
 [github_app]
-pr_commands = ""
+handle_pr_actions = []
 ```
 ##### GitHub app automatic tools for new code (PR push)
 In addition the running automatic tools when a PR is opened, the GitHub app can also respond to new code that is pushed to an open PR.
-Note that a local `.pr_agent.toml` file enables you to edit and customize the default parameters of any tool, not just the ones that are run automatically.
+The configuration toggle `handle_push_trigger` can be used to enable this feature.  
 The configuration parameter `push_commands` defines the list of tools that will be **run automatically** when new code is pushed to the PR.
 ```
 [github_app]
 handle_push_trigger = true
 push_commands = [
    "/describe --pr_description.add_original_user_description=true --pr_description.keep_original_user_title=true",
    "/auto_review -i --pr_reviewer.remove_previous_review_comment=true",
 ]
 ```
 The means that when new code is pused to the PR, the PR-Agent will run the `describe` and incremental `auto_review` tools.  
 For the describe tool, the `add_original_user_description` and `keep_original_user_title` parameters will be set to true.  
 For the `auto_review` tool, it will run in incremental mode, and the `remove_previous_review_comment` parameter will be set to true.
 Much like the configurations for `pr_commands`, you can override the default tool paramteres by uploading a local configuration file to the root of your repo.
 #### Editing the prompts
 The prompts for the various PR-Agent tools are defined in the `pr_agent/settings` folder.
@ -159,21 +187,28 @@ user="""
 Note that the new prompt will need to generate an output compatible with the relevant [post-process function](./pr_agent/tools/pr_description.py#L137).
 ### Working with GitHub Action
-You can configure settings in GitHub action by adding environment variables under the env section in `.github/workflows/pr_agent.yml` file. Some examples:
+You can configure settings in GitHub action by adding environment variables under the env section in `.github/workflows/pr_agent.yml` file. 
 Specifically, start by setting the following environment variables:
 ```yaml
      env:
-        # ... previous environment values
+        OPENAI_KEY: ${{ secrets.OPENAI_KEY }} # Make sure to add your OpenAI key to your repo secrets
-        OPENAI.ORG: "<Your organization name under your OpenAI account>"
+        GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # Make sure to add your GitHub token to your repo secrets
-        PR_REVIEWER.REQUIRE_TESTS_REVIEW: "false" # Disable tests review
+        github_action.auto_review: "true" # enable\disable auto review
-        PR_CODE_SUGGESTIONS.NUM_CODE_SUGGESTIONS: 6 # Increase number of code suggestions
+        github_action.auto_describe: "true" # enable\disable auto describe
-        github_action.auto_review: "true" # Enable auto review
+        github_action.auto_improve: "false" # enable\disable auto improve
        github_action.auto_describe: "true" # Enable auto describe
        github_action.auto_improve: "false" # Disable auto improve
 ```
-specifically, `github_action.auto_review`, `github_action.auto_describe` and `github_action.auto_improve` are used to enable/disable automatic tools that run when a new PR is opened.
+`github_action.auto_review`, `github_action.auto_describe` and `github_action.auto_improve` are used to enable/disable automatic tools that run when a new PR is opened.
 If not set, the default option is that only the `review` tool will run automatically when a new PR is opened.
 Note that you can give additional config parameters by adding environment variables to `.github/workflows/pr_agent.yml`, or by using a `.pr_agent.toml` file in the root of your repo, similar to the GitHub App usage.
 For example, you can set an environment variable: `pr_description.add_original_user_description=false`, or add a `.pr_agent.toml` file with the following content:
 ```
 [pr_description]
 add_original_user_description = false
 ```
 ### Changing a model
 See [here](pr_agent/algo/__init__.py) for the list of available models.
--- a/docs/DESCRIBE.md
+++ b/docs/DESCRIBE.md
@ -26,7 +26,7 @@ Under the section 'pr_description', the [configuration file](./../pr_agent/setti
 - `keep_original_user_title`: if set to true, the tool will keep the original PR title, and won't change it. Default is false.
 - `extra_instructions`: Optional extra instructions to the tool. For example: "focus on the changes in the file X. Ignore change in ...".
-
+- To enable `custom labels`, apply the configuration changes described [here](./GENERATE_CUSTOM_LABELS.md#configuration-changes)
 ### Markers template
 markers enable to easily integrate user's content and auto-generated content, with a template-like mechanism.
--- a/docs/GENERATE_CUSTOM_LABELS.md
+++ b/docs/GENERATE_CUSTOM_LABELS.md
@ -1,5 +1,5 @@
 # Generate Custom Labels
-The `generte_labels` tool scans the PR code changes, and given a list of labels and their descriptions, it automatically suggests labels that match the PR code changes.
+The `generate_labels` tool scans the PR code changes, and given a list of labels and their descriptions, it automatically suggests labels that match the PR code changes.
 It can be invoked manually by commenting on any PR:
 ```
@ -10,26 +10,31 @@ For example:
 If we wish to add detect changes to SQL queries in a given PR, we can add the following custom label along with its description:
 <kbd><img src=./../pics/custom_labels_list.png width="768"></kbd>
-When running the `generte_labels` tool on a PR that includes changes in SQL queries, it will automatically suggest the custom label:
+When running the `generate_labels` tool on a PR that includes changes in SQL queries, it will automatically suggest the custom label:
 <kbd><img src=./../pics/custom_label_published.png width="768"></kbd>
-### Configuration options
+### How to enable custom labels
-To enable custom labels, you need to add the following configuration to the [custom_labels file](./../pr_agent/settings/custom_labels.toml):
+
 Note that in addition to the dedicated tool `generate_labels`, the custom labels will also be used by the `review` and `describe` tools.
 #### CLI
 To enable custom labels, you need to apply the [configuration changes](#configuration-changes) to the [custom_labels file](./../pr_agent/settings/custom_labels.toml):
 #### GitHub Action and GitHub App
 To enable custom labels, you need to apply the [configuration changes](#configuration-changes) to the local `.pr_agent.toml` file in you repository.
 #### Configuration changes
 - Change `enable_custom_labels` to True: This will turn off the default labels and enable the custom labels provided in the custom_labels.toml file.
- - Add the custom labels to the custom_labels.toml file. It should be formatted as follows:
+ - Add the custom labels. It should be formatted as follows:
 ```
 [config]
 enable_custom_labels=true
 [custom_labels."Custom Label Name"]
 description = "Description of when AI should suggest this label"
 ```
 - You can add modify the list to include all the custom labels you wish to use in your repository.
-#### Github Action
+[custom_labels."Custom Label 2"]
-To use the `generte_labels` tool with Github Action:
+description = "Description of when AI should suggest this label 2"
 ```
 - Add the following file to your repository under `env` section in `.github/workflows/pr_agent.yml`
 - Comma separated list of custom labels and their descriptions
 - The number of labels and descriptions should be the same and in the same order (empty descriptions are allowed):
 ```
 CUSTOM_LABELS: "label1, label2, ..."
 CUSTOM_LABELS_DESCRIPTION: "label1 description, label2 description, ..."
 ```
--- a/docs/REVIEW.md
+++ b/docs/REVIEW.md
@ -25,6 +25,7 @@ Under the section 'pr_reviewer', the [configuration file](./../pr_agent/settings
 - `inline_code_comments`: if set to true, the tool will publish the code suggestions as comments on the code diff. Default is false.
 - `automatic_review`: if set to false, no automatic reviews will be done. Default is true.
 - `extra_instructions`: Optional extra instructions to the tool. For example: "focus on the changes in the file X. Ignore change in ...".
 - To enable `custom labels`, apply the configuration changes described [here](./GENERATE_CUSTOM_LABELS.md#configuration-changes) 
 ####  Incremental Mode
 For an incremental review, which only considers changes since the last PR-Agent review, this can be useful when working on the PR in an iterative manner, and you want to focus on the changes since the last review instead of reviewing the entire PR again, the following command can be used:
 ```
--- a/docs/TOOLS_GUIDE.md
+++ b/docs/TOOLS_GUIDE.md
@ -6,5 +6,6 @@
 - [SIMILAR_ISSUE](./SIMILAR_ISSUE.md)
 - [UPDATE CHANGELOG](./UPDATE_CHANGELOG.md)
 - [ADD DOCUMENTATION](./ADD_DOCUMENTATION.md)
 - [GENERATE CUSTOM LABELS](./GENERATE_CUSTOM_LABELS.md)
 See the **[installation guide](/INSTALL.md)** for instructions on how to setup PR-Agent.
--- a/pr_agent/algo/utils.py
+++ b/pr_agent/algo/utils.py
@ -101,7 +101,8 @@ def parse_code_suggestion(code_suggestions: dict, gfm_supported: bool=True) -> s
                markdown_text += f"   **{sub_key}:** {sub_value}\n"
            if not gfm_supported:
                if "relevant line" not in sub_key.lower(): # nicer presentation
-                        markdown_text = markdown_text.rstrip('\n') + "\\\n"
+                        # markdown_text = markdown_text.rstrip('\n') + "\\\n" # works for gitlab
                        markdown_text = markdown_text.rstrip('\n') + "   \n"  # works for gitlab and bitbucker
    markdown_text += "\n"
    return markdown_text
@ -307,6 +308,9 @@ def try_fix_yaml(review_text: str) -> dict:
 def set_custom_labels(variables):
    if not get_settings().config.enable_custom_labels:
        return
    labels = get_settings().custom_labels
    if not labels:
        # set default labels
--- a/pr_agent/git_providers/bitbucket_provider.py
+++ b/pr_agent/git_providers/bitbucket_provider.py
@ -142,10 +142,15 @@ class BitbucketProvider(GitProvider):
    def remove_initial_comment(self):
        try:
            for comment in self.temp_comments:
-                self.pr.delete(f"comments/{comment}")
+                self.remove_comment(comment)
        except Exception as e:
            get_logger().exception(f"Failed to remove temp comments, error: {e}")
    def remove_comment(self, comment):
        try:
            self.pr.delete(f"comments/{comment}")
        except Exception as e:
            get_logger().exception(f"Failed to remove comment, error: {e}")
    # funtion to create_inline_comment
    def create_inline_comment(self, body: str, relevant_file: str, relevant_line_in_file: str):
--- a/pr_agent/git_providers/codecommit_provider.py
+++ b/pr_agent/git_providers/codecommit_provider.py
@ -221,6 +221,9 @@ class CodeCommitProvider(GitProvider):
    def remove_initial_comment(self):
        return ""  # not implemented yet
    def remove_comment(self, comment):
        return ""  # not implemented yet
    def publish_inline_comment(self, body: str, relevant_file: str, relevant_line_in_file: str):
        # https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/codecommit/client/post_comment_for_compared_commit.html
        raise NotImplementedError("CodeCommit provider does not support publishing inline comments yet")
--- a/pr_agent/git_providers/gerrit_provider.py
+++ b/pr_agent/git_providers/gerrit_provider.py
@ -396,5 +396,8 @@ class GerritProvider(GitProvider):
        # shutil.rmtree(self.repo_path)
        pass
    def remove_comment(self, comment):
        pass
    def get_pr_branch(self):
        return self.repo.head
--- a/pr_agent/git_providers/git_provider.py
+++ b/pr_agent/git_providers/git_provider.py
@ -71,6 +71,10 @@ class GitProvider(ABC):
    def remove_initial_comment(self):
        pass
    @abstractmethod
    def remove_comment(self, comment):
        pass
    @abstractmethod
    def get_languages(self):
        pass
--- a/pr_agent/git_providers/github_provider.py
+++ b/pr_agent/git_providers/github_provider.py
@ -50,7 +50,7 @@ class GithubProvider(GitProvider):
    def get_incremental_commits(self):
        self.commits = list(self.pr.get_commits())
-        self.get_previous_review()
+        self.previous_review = self.get_previous_review(full=True, incremental=True)
        if self.previous_review:
            self.incremental.commits_range = self.get_commit_range()
            # Get all files changed during the commit range
@ -63,7 +63,7 @@ class GithubProvider(GitProvider):
    def get_commit_range(self):
        last_review_time = self.previous_review.created_at
-        first_new_commit_index = 0
+        first_new_commit_index = None
        for index in range(len(self.commits) - 1, -1, -1):
            if self.commits[index].commit.author.date > last_review_time:
                self.incremental.first_new_commit_sha = self.commits[index].sha
@ -71,15 +71,21 @@ class GithubProvider(GitProvider):
            else:
                self.incremental.last_seen_commit_sha = self.commits[index].sha
                break
-        return self.commits[first_new_commit_index:]
+        return self.commits[first_new_commit_index:] if first_new_commit_index is not None else []
-    def get_previous_review(self):
+    def get_previous_review(self, *, full: bool, incremental: bool):
-        self.previous_review = None
+        if not (full or incremental):
            raise ValueError("At least one of full or incremental must be True")
        if not getattr(self, "comments", None):
            self.comments = list(self.pr.get_issue_comments())
        prefixes = []
        if full:
            prefixes.append("## PR Analysis")
        if incremental:
            prefixes.append("## Incremental PR Review")
        for index in range(len(self.comments) - 1, -1, -1):
-            if self.comments[index].body.startswith("## PR Analysis") or self.comments[index].body.startswith("## Incremental PR Review"):
+            if any(self.comments[index].body.startswith(prefix) for prefix in prefixes):
-                self.previous_review = self.comments[index]
+                return self.comments[index]
                break
    def get_files(self):
        if self.incremental.is_incremental and self.file_set:
@ -218,10 +224,16 @@ class GithubProvider(GitProvider):
        try:
            for comment in getattr(self.pr, 'comments_list', []):
                if comment.is_temporary:
-                    comment.delete()
+                    self.remove_comment(comment)
        except Exception as e:
            get_logger().exception(f"Failed to remove initial comment, error: {e}")
    def remove_comment(self, comment):
        try:
            comment.delete()
        except Exception as e:
            get_logger().exception(f"Failed to remove comment, error: {e}")
    def get_title(self):
        return self.pr.title
--- a/pr_agent/git_providers/gitlab_provider.py
+++ b/pr_agent/git_providers/gitlab_provider.py
@ -287,10 +287,16 @@ class GitLabProvider(GitProvider):
    def remove_initial_comment(self):
        try:
            for comment in self.temp_comments:
-                comment.delete()
+                self.remove_comment(comment)
        except Exception as e:
            get_logger().exception(f"Failed to remove temp comments, error: {e}")
    def remove_comment(self, comment):
        try:
            comment.delete()
        except Exception as e:
            get_logger().exception(f"Failed to remove comment, error: {e}")
    def get_title(self):
        return self.mr.title
--- a/pr_agent/git_providers/local_git_provider.py
+++ b/pr_agent/git_providers/local_git_provider.py
@ -140,6 +140,9 @@ class LocalGitProvider(GitProvider):
    def remove_initial_comment(self):
        pass  # Not applicable to the local git provider, but required by the interface
    def remove_comment(self, comment):
        pass  # Not applicable to the local git provider, but required by the interface
    def get_languages(self):
        """
        Calculate percentage of languages in repository. Used for hunk prioritisation.
--- a/pr_agent/servers/github_action_runner.py
+++ b/pr_agent/servers/github_action_runner.py
@ -19,10 +19,6 @@ async def run_action():
    OPENAI_KEY = os.environ.get('OPENAI_KEY') or os.environ.get('OPENAI.KEY')
    OPENAI_ORG = os.environ.get('OPENAI_ORG') or os.environ.get('OPENAI.ORG')
    GITHUB_TOKEN = os.environ.get('GITHUB_TOKEN')
    CUSTOM_LABELS = os.environ.get('CUSTOM_LABELS')
    CUSTOM_LABELS_DESCRIPTIONS = os.environ.get('CUSTOM_LABELS_DESCRIPTIONS')
    # CUSTOM_LABELS is a comma separated list of labels (string), convert to list and strip spaces
    get_settings().set("CONFIG.PUBLISH_OUTPUT_PROGRESS", False)
    # Check if required environment variables are set
@ -38,7 +34,6 @@ async def run_action():
    if not GITHUB_TOKEN:
        print("GITHUB_TOKEN not set")
        return
    # CUSTOM_LABELS_DICT = handle_custom_labels(CUSTOM_LABELS, CUSTOM_LABELS_DESCRIPTIONS)
    # Set the environment variables in the settings
    get_settings().set("OPENAI.KEY", OPENAI_KEY)
@ -46,7 +41,6 @@ async def run_action():
        get_settings().set("OPENAI.ORG", OPENAI_ORG)
    get_settings().set("GITHUB.USER_TOKEN", GITHUB_TOKEN)
    get_settings().set("GITHUB.DEPLOYMENT_TYPE", "user")
    # get_settings().set("CUSTOM_LABELS", CUSTOM_LABELS_DICT)
    # Load the event payload
    try:
@ -104,31 +98,5 @@ async def run_action():
                        await PRAgent().handle_request(url, body)
 def handle_custom_labels(CUSTOM_LABELS, CUSTOM_LABELS_DESCRIPTIONS):
    if CUSTOM_LABELS:
        CUSTOM_LABELS = [x.strip() for x in CUSTOM_LABELS.split(',')]
    else:
        # Set default labels
        CUSTOM_LABELS = ['Bug fix', 'Tests', 'Bug fix with tests', 'Refactoring', 'Enhancement', 'Documentation',
                         'Other']
        print(f"Using default labels: {CUSTOM_LABELS}")
    if CUSTOM_LABELS_DESCRIPTIONS:
        CUSTOM_LABELS_DESCRIPTIONS = [x.strip() for x in CUSTOM_LABELS_DESCRIPTIONS.split(',')]
    else:
        # Set default labels
        CUSTOM_LABELS_DESCRIPTIONS = ['Fixes a bug in the code', 'Adds or modifies tests',
                                     'Fixes a bug in the code and adds or modifies tests',
                                     'Refactors the code without changing its functionality',
                                     'Adds new features or functionality',
                                     'Adds or modifies documentation',
                                     'Other changes that do not fit in any of the above categories']
        print(f"Using default labels: {CUSTOM_LABELS_DESCRIPTIONS}")
    # create a dictionary of labels and descriptions
    CUSTOM_LABELS_DICT = dict()
    for i in range(len(CUSTOM_LABELS)):
        CUSTOM_LABELS_DICT[CUSTOM_LABELS[i]] = {'description': CUSTOM_LABELS_DESCRIPTIONS[i]}
    return CUSTOM_LABELS_DICT
 if __name__ == '__main__':
    asyncio.run(run_action())
--- a/pr_agent/servers/github_app.py
+++ b/pr_agent/servers/github_app.py
@ -1,7 +1,7 @@
 import copy
 import os
-import time
+import asyncio.locks
-from typing import Any, Dict
+from typing import Any, Dict, List, Tuple
 import uvicorn
 from fastapi import APIRouter, FastAPI, HTTPException, Request, Response
@ -14,8 +14,9 @@ from pr_agent.algo.utils import update_settings_from_args
 from pr_agent.config_loader import get_settings, global_settings
 from pr_agent.git_providers import get_git_provider
 from pr_agent.git_providers.utils import apply_repo_settings
 from pr_agent.git_providers.git_provider import IncrementalPR
 from pr_agent.log import LoggingFormat, get_logger, setup_logger
-from pr_agent.servers.utils import verify_signature
+from pr_agent.servers.utils import verify_signature, DefaultDictWithTimeout
 setup_logger(fmt=LoggingFormat.JSON)
@ -47,6 +48,7 @@ async def handle_marketplace_webhooks(request: Request, response: Response):
    body = await get_body(request)
    get_logger().info(f'Request body:\n{body}')
 async def get_body(request):
    try:
        body = await request.json()
@ -61,7 +63,9 @@ async def get_body(request):
    return body
-_duplicate_requests_cache = {}
+_duplicate_requests_cache = DefaultDictWithTimeout(ttl=get_settings().github_app.duplicate_requests_cache_ttl)
 _duplicate_push_triggers = DefaultDictWithTimeout(ttl=get_settings().github_app.push_trigger_pending_tasks_ttl)
 _pending_task_duplicate_push_conditions = DefaultDictWithTimeout(asyncio.locks.Condition, ttl=get_settings().github_app.push_trigger_pending_tasks_ttl)
 async def handle_request(body: Dict[str, Any], event: str):
@ -109,26 +113,99 @@ async def handle_request(body: Dict[str, Any], event: str):
    # handle pull_request event:
    #   automatically review opened/reopened/ready_for_review PRs as long as they're not in draft,
    #   as well as direct review requests from the bot
-    elif event == 'pull_request':
+    elif event == 'pull_request' and action != 'synchronize':
-        pull_request = body.get("pull_request")
+        pull_request, api_url = _check_pull_request_event(action, body, log_context, bot_user)
-        if not pull_request:
+        if not (pull_request and api_url):
            return {}
        api_url = pull_request.get("url")
        if not api_url:
            return {}
        log_context["api_url"] = api_url
        if pull_request.get("draft", True) or pull_request.get("state") != "open" or pull_request.get("user", {}).get("login", "") == bot_user:
            return {}
        if action in get_settings().github_app.handle_pr_actions:
            if action == "review_requested":
                if body.get("requested_reviewer", {}).get("login", "") != bot_user:
                    return {}
-                if pull_request.get("created_at") == pull_request.get("updated_at"):
+            get_logger().info(f"Performing review for {api_url=} because of {event=} and {action=}")
-                    # avoid double reviews when opening a PR for the first time
+            await _perform_commands(get_settings().github_app.pr_commands, agent, body, api_url, log_context)
    # handle pull_request event with synchronize action - "push trigger" for new commits
    elif event == 'pull_request' and action == 'synchronize' and get_settings().github_app.handle_push_trigger:
        pull_request, api_url = _check_pull_request_event(action, body, log_context, bot_user)
        if not (pull_request and api_url):
            return {}
-            get_logger().info(f"Performing review because of event={event} and action={action}")
+
        # TODO: do we still want to get the list of commits to filter bot/merge commits?
        before_sha = body.get("before")
        after_sha = body.get("after")
        merge_commit_sha = pull_request.get("merge_commit_sha")
        if before_sha == after_sha:
            return {}
        if get_settings().github_app.push_trigger_ignore_merge_commits and after_sha == merge_commit_sha:
            return {}
        if get_settings().github_app.push_trigger_ignore_bot_commits and body.get("sender", {}).get("login", "") == bot_user:
            return {}
        # Prevent triggering multiple times for subsequent push triggers when one is enough:
        # The first push will trigger the processing, and if there's a second push in the meanwhile it will wait.
        # Any more events will be discarded, because they will all trigger the exact same processing on the PR.
        # We let the second event wait instead of discarding it because while the first event was being processed,
        # more commits may have been pushed that led to the subsequent events,
        # so we keep just one waiting as a delegate to trigger the processing for the new commits when done waiting.
        current_active_tasks = _duplicate_push_triggers.setdefault(api_url, 0)
        max_active_tasks = 2 if get_settings().github_app.push_trigger_pending_tasks_backlog else 1
        if current_active_tasks < max_active_tasks:
            # first task can enter, and second tasks too if backlog is enabled
            get_logger().info(
                f"Continue processing push trigger for {api_url=} because there are {current_active_tasks} active tasks"
            )
            _duplicate_push_triggers[api_url] += 1
        else:
            get_logger().info(
                f"Skipping push trigger for {api_url=} because another event already triggered the same processing"
            )
            return {}
        async with _pending_task_duplicate_push_conditions[api_url]:
            if current_active_tasks == 1:
                # second task waits
                get_logger().info(
                    f"Waiting to process push trigger for {api_url=} because the first task is still in progress"
                )
                await _pending_task_duplicate_push_conditions[api_url].wait()
                get_logger().info(f"Finished waiting to process push trigger for {api_url=} - continue with flow")
        try:
            if get_settings().github_app.push_trigger_wait_for_initial_review and not get_git_provider()(api_url, incremental=IncrementalPR(True)).previous_review:
                get_logger().info(f"Skipping incremental review because there was no initial review for {api_url=} yet")
                return {}
            get_logger().info(f"Performing incremental review for {api_url=} because of {event=} and {action=}")
            await _perform_commands(get_settings().github_app.push_commands, agent, body, api_url, log_context)
        finally:
            # release the waiting task block
            async with _pending_task_duplicate_push_conditions[api_url]:
                _pending_task_duplicate_push_conditions[api_url].notify(1)
                _duplicate_push_triggers[api_url] -= 1
    get_logger().info("event or action does not require handling")
    return {}
 def _check_pull_request_event(action: str, body: dict, log_context: dict, bot_user: str) -> Tuple[Dict[str, Any], str]:
    invalid_result = {}, ""
    pull_request = body.get("pull_request")
    if not pull_request:
        return invalid_result
    api_url = pull_request.get("url")
    if not api_url:
        return invalid_result
    log_context["api_url"] = api_url
    if pull_request.get("draft", True) or pull_request.get("state") != "open" or pull_request.get("user", {}).get("login", "") == bot_user:
        return invalid_result
    if action in ("review_requested", "synchronize") and pull_request.get("created_at") == pull_request.get("updated_at"):
        # avoid double reviews when opening a PR for the first time
        return invalid_result
    return pull_request, api_url
 async def _perform_commands(commands: List[str], agent: PRAgent, body: dict, api_url: str, log_context: dict):
    apply_repo_settings(api_url)
-            for command in get_settings().github_app.pr_commands:
+    for command in commands:
        split_command = command.split(" ")
        command = split_command[0]
        args = split_command[1:]
@ -139,9 +216,6 @@ async def handle_request(body: Dict[str, Any], event: str):
        with get_logger().contextualize(**log_context):
            await agent.handle_request(api_url, new_command)
    get_logger().info("event or action does not require handling")
    return {}
 def _is_duplicate_request(body: Dict[str, Any]) -> bool:
    """
@ -150,13 +224,8 @@ def _is_duplicate_request(body: Dict[str, Any]) -> bool:
    """
    request_hash = hash(str(body))
    get_logger().info(f"request_hash: {request_hash}")
-    request_time = time.monotonic()
+    is_duplicate = _duplicate_requests_cache.get(request_hash, False)
-    ttl = get_settings().github_app.duplicate_requests_cache_ttl  # in seconds
+    _duplicate_requests_cache[request_hash] = True
    to_delete = [key for key, key_time in _duplicate_requests_cache.items() if request_time - key_time > ttl]
    for key in to_delete:
        del _duplicate_requests_cache[key]
    is_duplicate = request_hash in _duplicate_requests_cache
    _duplicate_requests_cache[request_hash] = request_time
    if is_duplicate:
        get_logger().info(f"Ignoring duplicate request {request_hash}")
    return is_duplicate
--- a/pr_agent/servers/utils.py
+++ b/pr_agent/servers/utils.py
@ -1,5 +1,8 @@
 import hashlib
 import hmac
 import time
 from collections import defaultdict
 from typing import Callable, Any
 from fastapi import HTTPException
@ -25,3 +28,59 @@ def verify_signature(payload_body, secret_token, signature_header):
 class RateLimitExceeded(Exception):
    """Raised when the git provider API rate limit has been exceeded."""
    pass
 class DefaultDictWithTimeout(defaultdict):
    """A defaultdict with a time-to-live (TTL)."""
    def __init__(
        self,
        default_factory: Callable[[], Any] = None,
        ttl: int = None,
        refresh_interval: int = 60,
        update_key_time_on_get: bool = True,
        *args,
        **kwargs,
    ):
        """
        Args:
            default_factory: The default factory to use for keys that are not in the dictionary.
            ttl: The time-to-live (TTL) in seconds.
            refresh_interval: How often to refresh the dict and delete items older than the TTL.
            update_key_time_on_get: Whether to update the access time of a key also on get (or only when set).
        """
        super().__init__(default_factory, *args, **kwargs)
        self.__key_times = dict()
        self.__ttl = ttl
        self.__refresh_interval = refresh_interval
        self.__update_key_time_on_get = update_key_time_on_get
        self.__last_refresh = self.__time() - self.__refresh_interval
    @staticmethod
    def __time():
        return time.monotonic()
    def __refresh(self):
        if self.__ttl is None:
            return
        request_time = self.__time()
        if request_time - self.__last_refresh > self.__refresh_interval:
            return
        to_delete = [key for key, key_time in self.__key_times.items() if request_time - key_time > self.__ttl]
        for key in to_delete:
            del self[key]
        self.__last_refresh = request_time
    def __getitem__(self, __key):
        if self.__update_key_time_on_get:
            self.__key_times[__key] = self.__time()
        self.__refresh()
        return super().__getitem__(__key)
    def __setitem__(self, __key, __value):
        self.__key_times[__key] = self.__time()
        return super().__setitem__(__key, __value)
    def __delitem__(self, __key):
        del self.__key_times[__key]
        return super().__delitem__(__key)
--- a/pr_agent/settings/configuration.toml
+++ b/pr_agent/settings/configuration.toml
@ -24,6 +24,7 @@ num_code_suggestions=4
 inline_code_comments = false
 ask_and_reflect=false
 automatic_review=true
 remove_previous_review_comment=false
 extra_instructions = ""
 [pr_description] # /describe #
@ -86,6 +87,27 @@ pr_commands = [
    "/describe --pr_description.add_original_user_description=true --pr_description.keep_original_user_title=true",
    "/auto_review",
 ]
 # settings for "pull_request" event with "synchronize" action - used to detect and handle push triggers for new commits
 handle_push_trigger = false
 push_trigger_ignore_bot_commits = true
 push_trigger_ignore_merge_commits = true
 push_trigger_wait_for_initial_review = true
 push_trigger_pending_tasks_backlog = true
 push_trigger_pending_tasks_ttl = 300
 push_commands = [
    "/describe --pr_description.add_original_user_description=true --pr_description.keep_original_user_title=true",
    """/auto_review -i \
       --pr_reviewer.require_focused_review=false \
       --pr_reviewer.require_score_review=false \
       --pr_reviewer.require_tests_review=false \
       --pr_reviewer.require_security_review=false \
       --pr_reviewer.require_estimate_effort_to_review=false \
       --pr_reviewer.num_code_suggestions=0 \
       --pr_reviewer.inline_code_comments=false \
       --pr_reviewer.remove_previous_review_comment=true \
       --pr_reviewer.extra_instructions='' \
    """
 ]
 [gitlab]
 # URL to the gitlab service
--- a/pr_agent/tools/pr_reviewer.py
+++ b/pr_agent/tools/pr_reviewer.py
@ -100,6 +100,9 @@ class PRReviewer:
            if self.is_auto and not get_settings().pr_reviewer.automatic_review:
                get_logger().info(f'Automatic review is disabled {self.pr_url}')
                return None
            if self.is_auto and self.incremental.is_incremental and not self.incremental.first_new_commit_sha:
                get_logger().info(f"Incremental review is enabled for {self.pr_url} but there are no new commits")
                return None
            get_logger().info(f'Reviewing PR: {self.pr_url} ...')
@ -115,7 +118,9 @@ class PRReviewer:
                get_logger().info('Pushing PR review...')
                self.git_provider.publish_comment(pr_comment)
                self.git_provider.remove_initial_comment()
-
+                previous_review_comment = self._get_previous_review_comment()
                if previous_review_comment:
                    self._remove_previous_review_comment(previous_review_comment)
                if get_settings().pr_reviewer.inline_code_comments:
                    get_logger().info('Pushing inline code comments...')
                    self._publish_inline_code_comments()
@ -231,9 +236,13 @@ class PRReviewer:
        if self.incremental.is_incremental:
            last_commit_url = f"{self.git_provider.get_pr_url()}/commits/" \
                              f"{self.git_provider.incremental.first_new_commit_sha}"
            last_commit_msg = self.incremental.commits_range[0].commit.message if self.incremental.commits_range else ""
            incremental_review_markdown_text = f"Starting from commit {last_commit_url}"
            if last_commit_msg:
                incremental_review_markdown_text += f"  \n_({last_commit_msg.splitlines(keepends=False)[0]})_"
            data = OrderedDict(data)
            data.update({'Incremental PR Review': {
-                "⏮️ Review for commits since previous PR-Agent review": f"Starting from commit {last_commit_url}"}})
+                "⏮️ Review for commits since previous PR-Agent review": incremental_review_markdown_text}})
            data.move_to_end('Incremental PR Review', last=False)
        markdown_text = convert_to_markdown(data, self.git_provider.is_supported("gfm_markdown"))
@ -314,3 +323,26 @@ class PRReviewer:
                    break
        return question_str, answer_str
    def _get_previous_review_comment(self):
        """
        Get the previous review comment if it exists.
        """
        try:
            if get_settings().pr_reviewer.remove_previous_review_comment and hasattr(self.git_provider, "get_previous_review"):
                return self.git_provider.get_previous_review(
                    full=not self.incremental.is_incremental,
                    incremental=self.incremental.is_incremental,
                )
        except Exception as e:
            get_logger().exception(f"Failed to get previous review comment, error: {e}")
    def _remove_previous_review_comment(self, comment):
        """
        Remove the previous review comment if it exists.
        """
        try:
            if get_settings().pr_reviewer.remove_previous_review_comment and comment:
                self.git_provider.remove_comment(comment)
        except Exception as e:
            get_logger().exception(f"Failed to remove previous review comment, error: {e}")
Author	SHA1	Message	Date
mrT23	b57ec301e8	Merge pull request #408 from Codium-ai/tr/final_fixes fixed review	2023-10-29 09:02:48 -07:00
mrT23	71da20ea7e	better link	2023-10-29 18:01:50 +02:00
mrT23	c895657310	fixed review	2023-10-29 17:59:46 +02:00
Ori Kotek	eda20ccca9	Merge pull request #407 from zmeir/patch-1 Update Usage.md with new GitHub App features	2023-10-29 16:54:59 +02:00
Zohar Meir	aed113cd79	Update Usage.md with new GitHub App features	2023-10-29 16:33:38 +02:00
mrT23	0ab07a46c6	Merge pull request #405 from Codium-ai/tr/final_fixes Final Fixes and Updates to PR Agent	2023-10-29 06:02:34 -07:00
mrT23	5f32e28933	generate_labels	2023-10-29 15:02:16 +02:00
mrT23	7538c4dd2f	generate_labels	2023-10-29 14:59:50 +02:00
mrT23	e3845283f8	release notes	2023-10-29 14:58:36 +02:00
mrT23	a85921d3c5	release notes	2023-10-29 14:49:35 +02:00
mrT23	27b64fbcaf	release notes	2023-10-29 14:47:46 +02:00
mrT23	8d50f2ae82	release notes	2023-10-29 14:43:45 +02:00
mrT23	e97a03f522	Merge remote-tracking branch 'origin/main' into tr/final_fixes	2023-10-29 14:38:33 +02:00
mrT23	2e3344b5b0	Merge pull request #406 from Codium-ai/hl/custom_labels Add documentation to custom labels	2023-10-29 05:38:11 -07:00
mrT23	e1b51eace7	release notes	2023-10-29 14:37:04 +02:00
Hussam.lawen	49e3d5ec5f	Add documentation	2023-10-29 13:58:01 +02:00
mrT23	afa78ed3fb	final fixes	2023-10-29 13:07:22 +02:00
mrT23	72d5e4748e	final fixes	2023-10-29 13:05:15 +02:00
Ori Kotek	61d3e1ebf4	Merge pull request #394 from zmeir/zmeir-external-push_trigger Added support for automatic review on push event	2023-10-29 13:04:33 +02:00
mrT23	055b5ea700	final fixes	2023-10-29 13:03:12 +02:00
Hussam.lawen	3434296792	Documentation	2023-10-29 13:02:07 +02:00
mrT23	ae375c2ff0	final fixes	2023-10-29 13:01:55 +02:00
Hussam.lawen	3d5efdf4f3	Merge commit '9a585de36461a6941cb77009e5ab5f4b568a1ff7' into hl/custom_labels	2023-10-29 13:01:53 +02:00
Hussam.lawen	e83747300d	Merge branch 'main' of github.com:Codium-ai/pr-agent into hl/custom_labels	2023-10-29 12:09:43 +02:00
Hussam.lawen	013a689b33	generate_labels fix	2023-10-29 10:43:04 +02:00
Zohar Meir	e6bea76eee	Typo	2023-10-26 17:07:16 +03:00
zmeir	414f2b6767	Fix incremental review if there are no new commits (would have performed a full review instead)	2023-10-26 16:49:55 +03:00
zmeir	6541575a0e	Refactor to use pull_request synchronize event	2023-10-26 16:49:54 +03:00
zmeir	02570ea797	Remove previous review comment on push event	2023-10-26 16:46:54 +03:00
zmeir	65bb70a1dd	Added support for automatic review on push event The new feature can be enabled via the new configuration `github_app.handle_push_event`. To avoid any unwanted side-effects, the current default of this configuration is set to `false`. The high level flow (assuming the configuration is enabled): 1. receive push event from GitHub 2. extract branch and commits from event 3. find PR url for branch (currently does not support PRs from forks) 4. perform configured commands (e.g. `/describe`, `/review -i`) The push event flow is guarded by a backlog queue so that multiple push events on the same branch won't trigger multiple duplicate runs of the PR-Agent commands. Example timeline: 1. push 1 - start handling event 2. push 2 - waiting to be handled while push 1 event is still running 3. push 3 - event is dropped since handling it and handling push 2 is the same, so it is redundant 4. push 1 finished being handled 5. push 2 awakens from wait and continues handling (potentially reviewing the commits of both push 2 and push 3) All of these options are configurable and can be enabled/disabled as per the user's desire. Additional minor changes in this PR: 1. Created `DefaultDictWithTimeout` utility class to avoid too much boilerplate code in managing caches for outdated triggers. 2. Guard against running increment review when there are no new commits. 3. Minor styling changes for incremented review text.	2023-10-25 11:15:23 +03:00