diff --git a/KataGo/katago b/KataGo/katago index cf204b4..6de6aeb 100755 Binary files a/KataGo/katago and b/KataGo/katago differ diff --git a/KataGo/katago-bs52.exe b/KataGo/katago-bs52.exe index 3ffc4b2..98e365c 100644 Binary files a/KataGo/katago-bs52.exe and b/KataGo/katago-bs52.exe differ diff --git a/KataGo/katago.exe b/KataGo/katago.exe index 834a12b..017492c 100644 Binary files a/KataGo/katago.exe and b/KataGo/katago.exe differ diff --git a/LICENSE b/LICENSE index 9c26c53..89e8964 100644 --- a/LICENSE +++ b/LICENSE @@ -5,7 +5,7 @@ For on related licenses for these binaries and libraries see https://github.com/ 2. Icons from www.flaticon.com, used with permission with the following attributions: - New game/Load game/Config icons: made by Freepik from www.flaticon.com - Save icon: made by Pixel perfect from www.flaticon.com -- Next/Previous icon: made by RoundIcons from www.flaticon.com - other next/previous icons are derived work. +- Next/Previous icons: made by/derived from ones made by RoundIcons from www.flaticon.com Aside from the above, the license for all other content in this repository is as follows: diff --git a/README.md b/README.md index a056203..bf08978 100644 --- a/README.md +++ b/README.md @@ -84,20 +84,18 @@ while stronger players can pay more attention to smaller mistakes. Available AIs, with strength indicating an estimate for the default settings, are: * **[9p+]** **Default** is full KataGo, above professional level. +* **[~1d?]** **ScoreLoss** is KataGo making moves with probability `~ e^(-strength * points lost)`. * **Balance** is KataGo occasionally making weaker moves, attempting to win by ~2 points. * **Jigo** is KataGo aggressively making weaker moves, attempting to win by 0.5 points. * **[~4d]** **Policy** uses the top move from the policy network (it's 'shape sense' without reading), should be around high dan level depending on the model used. There is a setting to increase variety in the opening, but otherwise it plays deterministically. -* **[~5k]**: **P:Weighted** picks a random move weighted by the policy, as long as it's above `lower_bound`. `weaken_fac` uses `policy^(1/weaken_fac)`, increasing the chance for weaker moves. +* **[~2k]**: **P:Weighted** picks a random move weighted by the policy, as long as it's above `lower_bound`. `weaken_fac` uses `policy^(1/weaken_fac)`, increasing the chance for weaker moves. * **[~5k]**: **P:Pick** picks `pick_n + pick_frac * ` moves at random, and play the best move among them. The setting `pick_override` determines the minimum value at which this process is bypassed to play the best move instead, preventing obvious blunders. This, along with 'Weighted' are probably the best choice for kyu players who want a chance of winning without playing the sillier bots below. Variants of this strategy include: - * **[~5k]**: **P:Local** will pick such moves biased towards the last move with probability related to `local_stddev`. - * **[~10k]**: **~P:Tenuki** is biased in the opposite way as P:Local, using the same setting. + * **[~2k]**: **P:Local** will pick such moves biased towards the last move with probability related to `local_stddev`. + * **[~10k]**: **P:Tenuki** is biased in the opposite way as P:Local, using the same setting. * **[~10k]**: **P:Influence** is biased towards 4th+ line moves, with every line below that dividing both the chance of considering the move and the policy value by `influence_weight`. Consider setting `pick_frac=1.0` to only affect the policy weight. * **[~10k]**: **P:Territory** is biased in the opposite way, towards 1-3rd line moves, using the same setting. -* * **[~5k]**: **P:Noise** mixes the policy with `noise_strength` Dirichlet noise. At `noise_strength=0.9` play is near-random, while `noise_strength=0.7` is still quite strong. A threshold setting is included to avoid senseless first-line moves. - -Selecting the AI as either white or black opens up the option to configure it under 'Configure AI'. ### Analysis @@ -146,13 +144,14 @@ If you ever need to reset to the original settings, simply re-download the `conf ### Settings Panel * engine settings - * max_visits: the number of visits used in analyses and AI moves, higher is more accurate but slower. - * max_time: maximal time in seconds for analyses, even when the target number of visits has not been reached. - * fast_visits: the number of visits used for certain operations with fewer visits. - * katago: path to your KataGo executable. - * model: path to your KataGo model file. - * config: path to your KataGo config file. - * threads: number of threads to use in the KataGo analysis engine. + * max_visits: The number of visits used in analyses and AI moves, higher is more accurate but slower. + * max_time: Maximal time in seconds for analyses, even when the target number of visits has not been reached. + * fast_visits: The number of visits used for certain operations with fewer visits. + * wide_root_noise: Consider a wider variety of moves, using KataGo's `analysisWideRootNoise` option. Will affect both analysis and AIs such as ScoreLoss. (KataGo 1.4+ only, keep at 0.0 otherwise) + * katago: Path to your KataGo executable. + * model: Path to your KataGo model file. Note that the default model file included is an older 15 block one. Replace it with a new model from [here](https://github.com/lightvector/KataGo/releases) for maximal strength. + * config: Path to your KataGo config file. + * threads: Number of threads to use in the KataGo analysis engine. * game settings * init_size: the initial size of the board, on start-up. * init_komi: likewise, for komi. diff --git a/bots/ai2gtp.py b/bots/ai2gtp.py index 297f2e2..0417ac6 100644 --- a/bots/ai2gtp.py +++ b/bots/ai2gtp.py @@ -40,14 +40,15 @@ ENGINE_SETTINGS = { "threads": 1, } - engine = KataGoEngine(logger, ENGINE_SETTINGS) with open("config.json") as f: settings = json.load(f) all_ai_settings = settings["ai"] -all_ai_settings["dev"] = all_ai_settings["P:Noise"] +if bot == "dev": + engine.override_settings["maxVisits"] = 500 +all_ai_settings["dev"] = all_ai_settings["ScoreLoss"] ai_strategy = bot_strategy_names[bot] ai_settings = all_ai_settings[ai_strategy] @@ -110,15 +111,22 @@ while True: while len(handicaps) < min(n, bx * by): # really obscure cases handicaps.add(Move((random.randint(0, bx - 1), random.randint(0, by - 1)), player="B").sgf(board_size=game.board_size)) game.root.set_property("AB", list(handicaps)) + game._calculate_groups() gtp = [Move.from_sgf(m, game.board_size, "B").gtp() for m in handicaps] logger.log(f"Chose handicap placements as {gtp}", OUTPUT_ERROR) print(f"= {' '.join(gtp)}\n") sys.stdout.flush() game.analyze_all_nodes() # re-evaluate root + while engine.queries: # and make sure this gets processed + time.sleep(0.001) continue elif "set_free_handicap" in line: _, *stones = line.split(" ") game.root.set_property("AB", [Move.from_gtp(move.upper()).sgf(game.board_size) for move in stones]) + game._calculate_groups() + game.analyze_all_nodes() # re-evaluate root + while engine.queries: # and make sure this gets processed + time.sleep(0.001) logger.log(f"Set handicap placements to {game.root.get_list_property('AB')}", OUTPUT_ERROR) elif "genmove" in line: _, player = line.strip().split(" ") @@ -142,12 +150,7 @@ while True: move = game.play(Move(None, player=game.next_player)).move else: move, node = ai_move(game, ai_strategy, ai_settings) - if node is None: - while node is None: - logger.log(f"ERROR generating move, backing up with weighted.", OUTPUT_ERROR) - move, node = ai_move(game, "p:weighted", {"pick_override": 1.0, "lower_bound": 0.001, "weaken_fac": 1}) - else: - logger.log(f"Generated move {move}", OUTPUT_ERROR) + logger.log(f"Generated move {move}", OUTPUT_ERROR) print(f"= {move.gtp()}\n") sys.stdout.flush() malkovich_analysis(game.current_node) diff --git a/bots/ai_performance.pickle b/bots/ai_performance.pickle index 282eae6..b60c892 100644 Binary files a/bots/ai_performance.pickle and b/bots/ai_performance.pickle differ diff --git a/bots/engine_server.py b/bots/engine_server.py index b71d0cf..b008476 100644 --- a/bots/engine_server.py +++ b/bots/engine_server.py @@ -13,6 +13,7 @@ PORT = int(sys.argv[1]) if len(sys.argv) > 1 else 8587 ENGINE_SETTINGS = { "katago": "my/katago25", + # "katago": "KataGo/katago", "model": "KataGo/models/b15-1.3.2.txt.gz", "config": "KataGo/analysis_config.cfg", "max_visits": 50, diff --git a/bots/selfplay.py b/bots/selfplay.py index 315d9b1..aad1e6e 100644 --- a/bots/selfplay.py +++ b/bots/selfplay.py @@ -34,7 +34,7 @@ with open("config.json") as f: class AI: DEFAULT_ENGINE_SETTINGS = { - "katago": "KataGo/katago-bs", + "katago": "KataGo/katago", "model": "KataGo/models/b15-1.3.2.txt.gz", "config": "KataGo/analysis_config.cfg", "max_visits": 1, @@ -97,6 +97,8 @@ def retrieve_ais(selected_ais): test_ais = [ # AI("Jigo", {}, {"max_visits": 100}), + AI("Policy", {}, {"model": "my/model.bin.gz"}), + AI("Policy", {}, {"model": "KataGo/models/b10-1.3.txt.gz"}), AI("Policy", {}), AI("P:Local", {}), AI("P:Pick", {}), @@ -105,46 +107,46 @@ test_ais = [ AI("P:Local", {}), AI("P:Influence", {}), AI("P:Territory", {}), + AI("P:Weighted", {}), ] + for ai in test_ais: add_ai(ai) -N_GAMES = 20 +N_GAMES = 5 +BOARDSIZE = 19 ais_to_test = retrieve_ais(test_ais) results = defaultdict(list) -def play_games(black: AI, white: AI, n: int = N_GAMES): +def play_games(black: AI, white: AI): players = {"B": black, "W": white} engines = {"B": black.get_engine(), "W": white.get_engine()} tag = f"{black.name} vs {white.name}" try: - for i in range(n): - game = Game(Logger(), engines, {}) - game.root.add_list_property("PW", [white.name]) - game.root.add_list_property("PB", [black.name]) - start_time = time.time() - while not game.ended: - p = game.current_node.next_player - move = ai_move(game, players[p].strategy, players[p].ai_settings) - while not game.current_node.analysis_ready: - time.sleep(0.001) - game.game_id += f"_{game.current_node.format_score()}" - print(f"{tag}\tGame {i+1} finished in {time.time()-start_time:.1f}s {game.current_node.format_score()} -> {game.write_sgf('sgf_selfplay/')}", file=sys.stderr) - score = game.current_node.score - if score > 0.3: - black.elo_comp.beat(white.elo_comp) - elif score > -0.3: - black.elo_comp.tied(white.elo_comp) + game = Game(Logger(), engines, {"init_size": BOARDSIZE}) + game.root.add_list_property("PW", [white.name]) + game.root.add_list_property("PB", [black.name]) + start_time = time.time() + while not game.ended: + p = game.current_node.next_player + move = ai_move(game, players[p].strategy, players[p].ai_settings) + while not game.current_node.analysis_ready: + time.sleep(0.001) + game.game_id += f"_{game.current_node.format_score()}" + print(f"{tag}\tGame finished in {time.time()-start_time:.1f}s {game.current_node.format_score()} -> {game.write_sgf('sgf_selfplay/')}", file=sys.stderr) + score = game.current_node.score + if score > 0.3: + black.elo_comp.beat(white.elo_comp) + elif score > -0.3: + black.elo_comp.tied(white.elo_comp) - results[tag].append(score) - all_results.append((black.name, white.name, score)) + results[tag].append(score) + all_results.append((black.name, white.name, score)) - with open("bots/tmp.pickle", "wb") as f: - pickle.dump((ai_database, all_results), f) except Exception as e: print(f"Exception in playing {tag}: {e}") print(f"Exception in playing {tag}: {e}", file=sys.stderr) @@ -159,32 +161,36 @@ def fmt_score(score): print(len(ais_to_test), "ais to test") global_start = time.time() -with ThreadPoolExecutor(max_workers=16) as threadpool: - for b in ais_to_test: - for w in ais_to_test: - if b is not w: - threadpool.submit(play_games, b, w) +for n in range(N_GAMES): + for _, e in AI.ENGINES: # no caching/replays + e.shutdown() + AI.ENGINES = [] -print("POOL EXIT") + with ThreadPoolExecutor(max_workers=16) as threadpool: + for b in ais_to_test: + for w in ais_to_test: + if b is not w: + threadpool.submit(play_games, b, w) + print("POOL EXIT") -print("---- RESULTS ----") -for k, v in results.items(): - b_win = sum([s > 0.3 for s in v]) - w_win = sum([s < -0.3 for s in v]) - print(f"{b_win} {k} {w_win} : {list(map(fmt_score,v))}") + print(f"---- RESULTS ({n}) ----") + for k, v in results.items(): + b_win = sum([s > 0.3 for s in v]) + w_win = sum([s < -0.3 for s in v]) + print(f"{b_win} {k} {w_win} : {list(map(fmt_score,v))}") -print("---- ELO ----") -for ai in sorted(ai_database, key=lambda a: -a.elo_comp.rating): - wins = [(b, w, s) for (b, w, s) in all_results if s > 0.3 and b == ai.name or w == ai.name and s < -0.3] - losses = [(b, w, s) for (b, w, s) in all_results if s < -0.3 and b == ai.name or w == ai.name and s > -0.3] - draws = [(b, w, s) for (b, w, s) in all_results if -0.3 <= s <= 0.3 and (b == ai.name or w == ai.name)] - out = f"{'*' if ai in ais_to_test else ' '} {ai.name}: ELO {ai.elo_comp.rating:.1f} WINS {len(wins)} LOSSES {len(losses)} DRAWS {len(draws)}" - # print("Wins:",wins) - print(out) - print(out, file=sys.stderr) + print("---- ELO ----") + for ai in sorted(ai_database, key=lambda a: -a.elo_comp.rating): + wins = [(b, w, s) for (b, w, s) in all_results if s > 0.3 and b == ai.name or w == ai.name and s < -0.3] + losses = [(b, w, s) for (b, w, s) in all_results if s < -0.3 and b == ai.name or w == ai.name and s > -0.3] + draws = [(b, w, s) for (b, w, s) in all_results if -0.3 <= s <= 0.3 and (b == ai.name or w == ai.name)] + out = f"{'*' if ai in ais_to_test else ' '} {ai.name}: ELO {ai.elo_comp.rating:.1f} WINS {len(wins)} LOSSES {len(losses)} DRAWS {len(draws)}" + # print("Wins:",wins) + print(out) + print(out, file=sys.stderr) -with open(DB_FILENAME, "wb") as f: - pickle.dump((ai_database, all_results), f) + with open(DB_FILENAME, "wb") as f: + pickle.dump((ai_database, all_results), f) + print(f"Saving {len(all_results)} to pickle", file=sys.stderr) -print(f"Done! saving {len(all_results)} to pickle", file=sys.stderr) -print(f"Time taken {time.time()-global_start:.1f}s", file=sys.stderr) +print(f"Done!Time taken {time.time()-global_start:.1f}s", file=sys.stderr) diff --git a/bots/settings.py b/bots/settings.py index 0db6b04..1dd6282 100644 --- a/bots/settings.py +++ b/bots/settings.py @@ -1,5 +1,6 @@ bot_strategy_names = { - "dev": "P:Noise", + # "dev": "P:Noise", + "dev": "ScoreLoss", "dev-beta": "P:Weighted", "strong": "Policy", "influence": "P:Influence", @@ -12,7 +13,8 @@ bot_strategy_names = { greetings = { - "dev": "Policy+Dirichlet noise.", + # "dev": "Policy+Dirichlet noise.", + "dev": "Point loss-weighted random move.", "dev-beta": "Play a policy-weighted move.", "strong": "Play top policy move.", "influence": "Play an influential style.", diff --git a/config.json b/config.json index 0d4a0dd..a081305 100644 --- a/config.json +++ b/config.json @@ -7,6 +7,7 @@ "max_visits": 500, "fast_visits": 50, "max_time": 3.0, + "wide_root_noise": 0.0, "_enable_ownership": true }, "sgf": { @@ -64,6 +65,11 @@ "_help_left": "Will try to win by `target_score`, without further restrictions.", "_help_right": "Also affected by engine settings such as `max_visits`." }, + "ScoreLoss": { + "strength": 0.5, + "_help_left": "Plays moves weighted inversely by point loss.", + "_help_right": "Also affected by engine settings such as `max_visits`, likely to play more varied/weaker with higher visits." + }, "Policy": { "opening_moves": 0.05, "_help_left": "Strength is mainly affected by `model` in engine settings, but should be high dan regardless.", @@ -76,13 +82,6 @@ "lower_bound": 0.001, "weaken_fac": 1.25 }, - "P:Noise": { - "pick_override": 0.95, - "noise_strength": 0.6, - "lower_bound": 0.001, - "_help_left": "Adds `noise_strength` noise to the policy of all moved > 'lower_bound' and plays the top move.", - "_help_right": "Plays top move if policy value is above `pick_override` to avoid obvious mistakes. Noise above 0.9 is near random, below 0.7 is fairly strong." - }, "P:Pick": { "pick_override": 0.95, "pick_n": 5, @@ -102,9 +101,10 @@ "pick_override": 0.85, "stddev": 7.5, "pick_n": 5, - "pick_frac": 0.7, + "pick_frac": 0.5, + "endgame": 0.45, "_help_left": "Samples `pick_n + pick_frac * ` away from the last move and plays the best one.", - "_help_right": "Increase `stddev` makes it prefer moves further away." + "_help_right": "Increase `stddev` makes it prefer moves further away. Stops tenukiing after the 'endgame' fraction of the board is filled." }, "P:Influence": { "pick_override": 0.95, @@ -112,8 +112,9 @@ "pick_frac": 0.4, "threshold": 3.5, "line_weight": 10, + "endgame": 0.4, "_help_left": "Samples `pick_n + pick_frac * ` and plays the best one, biased to above the `threshold` line.", - "_help_right": "Increase `line_weight` to penalize moves near the edge more." + "_help_right": "Increase `line_weight` to penalize moves near the edge more. Stops strategy after the 'endgame' fraction of the board is filled." }, "P:Territory": { "pick_override": 0.95, @@ -121,8 +122,9 @@ "pick_frac": 0.4, "threshold": 3.5, "line_weight": 2, + "endgame": 0.4, "_help_left": "Samples `pick_n + pick_frac * ` and plays the best one, biased to below the `threshold` line.", - "_help_right": "Increase `line_weight` to penalize moves closer to the center more." + "_help_right": "Increase `line_weight` to penalize moves closer to the center more. Stops strategy after the 'endgame' fraction of the board is filled." } }, "board_ui": { diff --git a/core/ai.py b/core/ai.py index 87c4d1c..c86d904 100644 --- a/core/ai.py +++ b/core/ai.py @@ -9,7 +9,7 @@ from core.engine import EngineDiedException from core.game import Game, GameNode, IllegalMoveException, Move -def weighted_selection_without_replacement(items: List[Tuple[float, float, int, int]], pick_n: int) -> List[Tuple[float, float, int, int]]: +def weighted_selection_without_replacement(items: List[Tuple], pick_n: int) -> List[Tuple]: """For a list of tuples where the second element is a weight, returns random items with those weights, without replacement.""" elt = [(math.log(random.random()) / item[1], item) for item in items] # magic return [e[1] for e in heapq.nlargest(pick_n, elt)] # NB fine if too small @@ -45,7 +45,7 @@ def ai_move(game: Game, ai_mode: str, ai_settings: Dict) -> Tuple[Move, GameNode ai_thoughts += f"Using policy based strategy, base top 5 moves are {fmt_moves(policy_moves[:5])}. " if "policy" in ai_mode and cn.depth <= int(ai_settings["opening_moves"] * (game.board_size[0] * game.board_size[1])): ai_mode = "p:weighted" - ai_thoughts += f"Switching to weighted strategy in the opening {int(ai_settings['opening_moves'] * (game.board_size[0]*game.board_size[1]))} moves." + ai_thoughts += f"Switching to weighted strategy in the opening {int(ai_settings['opening_moves'] * (game.board_size[0]*game.board_size[1]))} moves. " ai_settings = {"pick_override": 0.9, "weaken_fac": 1, "lower_bound": 0.02} if top_5_pass: aimove = top_policy_move @@ -57,9 +57,14 @@ def ai_move(game: Game, ai_mode: str, ai_settings: Dict) -> Tuple[Move, GameNode aimove = top_policy_move ai_thoughts += f"Top policy move has weight > {ai_settings['pick_override']:.1%}, so overriding other strategies." elif "weighted" in ai_mode: - lower_bound = max(0, ai_settings["lower_bound"]) + lower_bound = max(0, ai_settings["lower_bound"]) * 2 # compensate for first halving in loop weaken_fac = max(0.01, ai_settings["weaken_fac"]) - weighted_coords = [(policy_grid[y][x], policy_grid[y][x] ** (1 / weaken_fac), x, y) for x in range(size[0]) for y in range(size[1]) if policy_grid[y][x] > lower_bound] + weighted_coords = [] + while not weighted_coords and lower_bound > 1e-6: # fix edge case where no moves are > lb + lower_bound /= 2 + weighted_coords = [ + (policy_grid[y][x], policy_grid[y][x] ** (1 / weaken_fac), x, y) for x in range(size[0]) for y in range(size[1]) if policy_grid[y][x] > lower_bound + ] top = weighted_selection_without_replacement(weighted_coords, 1) if top: best = top[0] @@ -72,7 +77,7 @@ def ai_move(game: Game, ai_mode: str, ai_settings: Dict) -> Tuple[Move, GameNode ai_thoughts += f"Playing policy-weighted random move {aimove.gtp()} ({policy_value:.1%})" + ( " because no other moves were found." if not top else f" because strategy is weighted (lower bound={lower_bound:.2%}, num moves > lb={len(weighted_coords)})." ) - elif "noise" in ai_mode: + elif "noise" in ai_mode: # DEPRECATED noise_str = ai_settings["noise_strength"] lower_bound = max(0, ai_settings["lower_bound"]) selected_policy_moves = [(pol, mv) for pol, mv in policy_moves if not mv.is_pass if pol > lower_bound] @@ -90,13 +95,18 @@ def ai_move(game: Game, ai_mode: str, ai_settings: Dict) -> Tuple[Move, GameNode legal_policy_moves = [(pol, mv) for pol, mv in policy_moves if not mv.is_pass if pol > 0] n_moves = int(ai_settings["pick_frac"] * len(legal_policy_moves) + ai_settings["pick_n"]) if "influence" in ai_mode or "territory" in ai_mode: + thr_line = ai_settings["threshold"] - 1 # zero-based - if "influence" in ai_mode: - weight = lambda x, y: (1 / ai_settings["line_weight"]) ** (max(0, thr_line - min(size[0] - 1 - x, x)) + max(0, thr_line - min(size[1] - 1 - y, y))) + if cn.depth >= ai_settings["endgame"] * size[0] * size[1]: + weighted_coords = [(policy_grid[y][x], 1, x, y) for x in range(size[0]) for y in range(size[1]) if policy_grid[y][x] > 0] + ai_thoughts += f"Generated equal weights as move number >= {ai_settings['endgame'] * size[0] * size[1]}. " else: - weight = lambda x, y: (1 / ai_settings["line_weight"]) ** (max(0, min(size[0] - 1 - x, x, size[1] - 1 - y, y) - thr_line)) - weighted_coords = [(policy_grid[y][x] * weight(x, y), weight(x, y), x, y) for x in range(size[0]) for y in range(size[1]) if policy_grid[y][x] > 0] - ai_thoughts += f"Generated weights for {ai_mode} according to weight factor {ai_settings['line_weight']} and distance from {thr_line+1}th line. " + if "influence" in ai_mode: + weight = lambda x, y: (1 / ai_settings["line_weight"]) ** (max(0, thr_line - min(size[0] - 1 - x, x)) + max(0, thr_line - min(size[1] - 1 - y, y))) + else: + weight = lambda x, y: (1 / ai_settings["line_weight"]) ** (max(0, min(size[0] - 1 - x, x, size[1] - 1 - y, y) - thr_line)) + weighted_coords = [(policy_grid[y][x] * weight(x, y), weight(x, y), x, y) for x in range(size[0]) for y in range(size[1]) if policy_grid[y][x] > 0] + ai_thoughts += f"Generated weights for {ai_mode} according to weight factor {ai_settings['line_weight']} and distance from {thr_line+1}th line. " elif "local" in ai_mode or "tenuki" in ai_mode: var = ai_settings["stddev"] ** 2 if not cn.move or cn.move.coords is None: @@ -108,8 +118,12 @@ def ai_move(game: Game, ai_mode: str, ai_settings: Dict) -> Tuple[Move, GameNode (policy_grid[y][x], math.exp(-0.5 * ((x - mx) ** 2 + (y - my) ** 2) / var), x, y) for x in range(size[0]) for y in range(size[1]) if policy_grid[y][x] > 0 ] if "tenuki" in ai_mode: - weighted_coords = [(p, 1 - w, x, y) for p, w, x, y in weighted_coords] - ai_thoughts += f"Generated weights based on one minus gaussian with variance {var} around coordinates {mx},{my}. " + if cn.depth < ai_settings["endgame"] * size[0] * size[1]: + weighted_coords = [(p, 1 - w, x, y) for p, w, x, y in weighted_coords] + ai_thoughts += f"Generated weights based on one minus gaussian with variance {var} around coordinates {mx},{my}. " + else: + weighted_coords = [(p, 1, x, y) for p, w, x, y in weighted_coords] + ai_thoughts += f"Generated equal weights as move number >= {ai_settings['endgame'] * size[0] * size[1]}. " else: ai_thoughts += f"Generated weights based on gaussian with variance {var} around coordinates {mx},{my}. " elif "pick" in ai_mode: @@ -132,33 +146,40 @@ def ai_move(game: Game, ai_mode: str, ai_settings: Dict) -> Tuple[Move, GameNode raise ValueError(f"Unknown AI mode {ai_mode}") else: # Engine based move candidate_ai_moves = cn.candidate_moves - if "balance" in ai_mode and candidate_ai_moves[0]["move"] != "pass": # don't play suicidal to balance score - pass when it's best - sign = cn.player_sign(cn.next_player) - sel_moves = [ # top move, or anything not too bad, or anything that makes you still ahead - move - for i, move in enumerate(candidate_ai_moves) - if i == 0 - or move["visits"] >= ai_settings["min_visits"] - and (move["pointsLost"] < ai_settings["random_loss"] or move["pointsLost"] < ai_settings["max_loss"] and sign * move["scoreLead"] > ai_settings["target_score"]) - ] - aimove = Move.from_gtp(random.choice(sel_moves)["move"], player=cn.next_player) - ai_thoughts += f"Balance strategy selected moves {sel_moves} based on target score and max points lost, and randomly chose {aimove.gtp()}." - elif "jigo" in ai_mode and candidate_ai_moves[0]["move"] != "pass": - sign = cn.player_sign(cn.next_player) - jigo_move = min(candidate_ai_moves, key=lambda move: abs(sign * move["scoreLead"] - ai_settings["target_score"])) - aimove = Move.from_gtp(jigo_move["move"], player=cn.next_player) - ai_thoughts += f"Jigo strategy found candidate moves {candidate_ai_moves} moves and chose {aimove.gtp()} as closest to 0.5 point win" + top_cand = Move.from_gtp(candidate_ai_moves[0]["move"], player=cn.next_player) + if top_cand.is_pass: # don't play suicidal to balance score - pass when it's best + aimove = top_cand + ai_thoughts += f"Top move is pass, so passing regardless of strategy." else: - if "default" not in ai_mode and "katago" not in ai_mode: - game.katrain.log(f"Unknown AI mode {ai_mode} or policy missing, using default.", OUTPUT_INFO) - ai_thoughts += f"Strategy {ai_mode} not found or unexpected fallback." - aimove = Move.from_gtp(candidate_ai_moves[0]["move"], player=cn.next_player) - ai_thoughts += f"Default strategy found {len(candidate_ai_moves)} moves returned from the engine and chose {aimove.gtp()} as top move" + if "balance" in ai_mode: + sign = cn.player_sign(cn.next_player) + sel_moves = [ # top move, or anything not too bad, or anything that makes you still ahead + move + for i, move in enumerate(candidate_ai_moves) + if i == 0 + or move["visits"] >= ai_settings["min_visits"] + and (move["pointsLost"] < ai_settings["random_loss"] or move["pointsLost"] < ai_settings["max_loss"] and sign * move["scoreLead"] > ai_settings["target_score"]) + ] + aimove = Move.from_gtp(random.choice(sel_moves)["move"], player=cn.next_player) + ai_thoughts += f"Balance strategy selected moves {sel_moves} based on target score and max points lost, and randomly chose {aimove.gtp()}." + elif "jigo" in ai_mode: + sign = cn.player_sign(cn.next_player) + jigo_move = min(candidate_ai_moves, key=lambda move: abs(sign * move["scoreLead"] - ai_settings["target_score"])) + aimove = Move.from_gtp(jigo_move["move"], player=cn.next_player) + ai_thoughts += f"Jigo strategy found {len(candidate_ai_moves)} candidate moves (best {top_cand.gtp()}) and chose {aimove.gtp()} as closest to 0.5 point win" + elif "scoreloss" in ai_mode: + c = ai_settings["strength"] + moves = [(d["pointsLost"], math.exp(-c * max(0, d["pointsLost"])), Move.from_gtp(d["move"], player=cn.next_player)) for d in candidate_ai_moves] + topmove = weighted_selection_without_replacement(moves, 1)[0] + aimove = topmove[2] + ai_thoughts += f"ScoreLoss strategy found {len(candidate_ai_moves)} candidate moves (best {top_cand.gtp()}) and chose {aimove.gtp()} (weight {topmove[1]:.3f}, point loss {topmove[0]:.1f}) based on score weights." + else: + if "default" not in ai_mode and "katago" not in ai_mode: + game.katrain.log(f"Unknown AI mode {ai_mode} or policy missing, using default.", OUTPUT_INFO) + ai_thoughts += f"Strategy {ai_mode} not found or unexpected fallback." + aimove = top_cand + ai_thoughts += f"Default strategy found {len(candidate_ai_moves)} moves returned from the engine and chose {aimove.gtp()} as top move" game.katrain.log(f"AI thoughts: {ai_thoughts}", OUTPUT_DEBUG) - try: - played_node = game.play(aimove) - played_node.ai_thoughts = ai_thoughts - return aimove, played_node - except IllegalMoveException as e: - game.katrain.log(f"AI Strategy {ai_mode} generated illegal move {aimove.gtp()}: {e}", OUTPUT_ERROR) - return None, None + played_node = game.play(aimove) + played_node.ai_thoughts = ai_thoughts + return aimove, played_node diff --git a/core/common.py b/core/common.py index 7427759..b667dde 100644 --- a/core/common.py +++ b/core/common.py @@ -1,6 +1,7 @@ from typing import Any, List, Tuple OUTPUT_ERROR = -1 +OUTPUT_KATAGO_STDERR = -0.5 OUTPUT_INFO = 0 OUTPUT_DEBUG = 1 OUTPUT_EXTRA_DEBUG = 2 diff --git a/core/engine.py b/core/engine.py index c662207..83a3d54 100644 --- a/core/engine.py +++ b/core/engine.py @@ -6,7 +6,7 @@ import threading import time from typing import Callable, Optional -from core.common import OUTPUT_DEBUG, OUTPUT_ERROR, OUTPUT_EXTRA_DEBUG +from core.common import OUTPUT_DEBUG, OUTPUT_ERROR, OUTPUT_EXTRA_DEBUG, OUTPUT_KATAGO_STDERR from core.game_node import GameNode @@ -35,14 +35,16 @@ class KataGoEngine: self.query_counter = 0 self.katago_process = None self.base_priority = 0 + self.override_settings = {} # mainly for bot scripts to hook into self._lock = threading.Lock() self.start() self.analysis_thread = threading.Thread(target=self._analysis_read_thread, daemon=True).start() + self.stderr_thread = threading.Thread(target=self._read_stderr_thread, daemon=True).start() def start(self): try: self.katrain.log(f"Starting KataGo with {self.command}", OUTPUT_DEBUG) - self.katago_process = subprocess.Popen(self.command, stdin=subprocess.PIPE, stdout=subprocess.PIPE) + self.katago_process = subprocess.Popen(self.command, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE) except FileNotFoundError as e: self.katrain.log( f"Starting kata with command '{self.command}' failed with error {e}. Please make sure the 'katago' value under 'engine' in settings points to the correct KataGo executable.", @@ -70,6 +72,15 @@ class KataGoEngine: def is_idle(self): return not self.queries + def _read_stderr_thread(self): + while self.katago_process is not None: + try: + line = self.katago_process.stderr.readline() + if line: + self.katrain.log(line.decode(), OUTPUT_KATAGO_STDERR) + except: + return + def _analysis_read_thread(self): while self.katago_process is not None: try: @@ -81,30 +92,27 @@ class KataGoEngine: if not line: continue analysis = json.loads(line) - if analysis["id"] in self.queries: - query_id = analysis["id"] - callback, error_callback, start_time, next_move = self.queries[query_id] - else: + if analysis["id"] not in self.queries: self.katrain.log(f"Query result {analysis['id']} discarded -- recent new game?", OUTPUT_DEBUG) continue + query_id = analysis["id"] + callback, error_callback, start_time, next_move = self.queries[query_id] + del self.queries[query_id] if "error" in analysis: if error_callback: error_callback(analysis) elif not (next_move and "Illegal move" in analysis["error"]): # sweep self.katrain.log(f"{analysis} received from KataGo", OUTPUT_ERROR) - continue else: - callback, error_callback, start_time, next_move = self.queries[query_id] time_taken = time.time() - start_time self.katrain.log(f"[{time_taken:.1f}][{analysis['id']}] KataGo Analysis Received: {analysis.keys()}", OUTPUT_DEBUG) self.katrain.log(line, OUTPUT_EXTRA_DEBUG) - del self.queries[query_id] try: callback(analysis) except Exception as e: self.katrain.log(f"Error in engine callback for query {query_id}: {e}", OUTPUT_ERROR) - if getattr(self.katrain, "update_state", None): # easier mocking etc - self.katrain.update_state() + if getattr(self.katrain, "update_state", None): # easier mocking etc + self.katrain.update_state() def send_query(self, query, callback, error_callback, next_move=None): with self._lock: @@ -144,6 +152,12 @@ class KataGoEngine: visits = self.config["fast_visits"] size_x, size_y = analysis_node.board_size + settings = self.override_settings + if time_limit: + settings["maxTime"] = self.config["max_time"] + if self.config.get("wide_root_noise",0.0) > 0.0: # don't send if 0.0, so older versions don't error + settings["wideRootNoise"] = self.config["wide_root_noise"] + query = { "rules": self.get_rules(analysis_node), "priority": self.base_priority + priority, @@ -155,7 +169,6 @@ class KataGoEngine: "includeOwnership": ownership, "includePolicy": not next_move, "moves": [[m.player, m.gtp()] for m in moves], - "overrideSettings": {"maxTime": self.config["max_time"] if time_limit else 1000.0} - # "overrideSettings": {"playoutDoublingAdvantage": 3.0, "playoutDoublingAdvantagePla": 'BLACK' if not moves or moves[-1].player == 'W' else "WHITE"} + "overrideSettings": settings } self.send_query(query, callback, error_callback, next_move) diff --git a/core/game.py b/core/game.py index d2eb1ea..c4a194d 100644 --- a/core/game.py +++ b/core/game.py @@ -233,14 +233,14 @@ class Game: return self.current_node.format_score(score) def __repr__(self): - return "\n".join("".join(Move.PLAYERS[self.chains[c][0].player] if c >= 0 else "-" for c in line) for line in self.board) + f"\ncaptures: {self.prisoner_count}" + return "\n".join("".join(self.chains[c][0].player if c >= 0 else "-" for c in line) for line in self.board) + f"\ncaptures: {self.prisoner_count}" def write_sgf(self, path=None, trainer_config={}, save_feedback=(True,), eval_thresholds=(0,)): black, white = self.root.get_property("PB"), self.root.get_property("PW") black = re.sub(r"['<>:\"/\\|?*]", "", black or "Black") white = re.sub(r"['<>:\"/\\|?*]", "", white or "White") game_name = f"katrain_{black} vs {white} {self.game_id}" - file_name = os.path.join(path, f"{game_name}.sgf") + file_name = os.path.abspath(os.path.join(path, f"{game_name}.sgf")) os.makedirs(os.path.dirname(file_name), exist_ok=True) show_dots_for = {p: trainer_config.get("eval_show_ai", True) or "ai" not in self.katrain.controls.player_mode(p) for p in Move.PLAYERS} diff --git a/core/game_node.py b/core/game_node.py index beebe4d..35d24a2 100644 --- a/core/game_node.py +++ b/core/game_node.py @@ -169,10 +169,9 @@ class GameNode(SGFNode): top_polmove = polmoves[0][1] if polmoves else Move(None) # if no info at all, pass return [{**self.analysis["root"], "pointsLost": 0, "order": 0, "move": top_polmove.gtp()}] # single visit -> go by policy/root - return sorted( - [{"pointsLost": self.player_sign(self.next_player) * (self.analysis["root"]["scoreLead"] - d["scoreLead"]), **d} for d in self.analysis["moves"].values()], - key=lambda d: (d["order"], d["pointsLost"]), - ) + root_score = self.analysis["root"]["scoreLead"] + move_dicts = list(self.analysis["moves"].values()) # prevent incoming analysis from causing crash + return sorted([{"pointsLost": self.player_sign(self.next_player) * (root_score - d["scoreLead"]), **d} for d in move_dicts], key=lambda d: (d["order"], d["pointsLost"])) @property def policy_ranking(self) -> Optional[List[Tuple[float, Move]]]: # return moves from highest policy value to lowest diff --git a/core/sgf_parser.py b/core/sgf_parser.py index 9e1a994..66c54b8 100644 --- a/core/sgf_parser.py +++ b/core/sgf_parser.py @@ -28,7 +28,7 @@ class Move: """Initialize a move from SGF coordinates and player""" if sgf_coords == "" or Move.SGF_COORD.index(sgf_coords[0]) == board_size[0]: # some servers use [tt] for pass return cls(coords=None, player=player) - return cls(coords=(Move.SGF_COORD.index(sgf_coords[0]), board_size[1] - Move.SGF_COORD.index(sgf_coords[1]) - 1), player=player,) + return cls(coords=(Move.SGF_COORD.index(sgf_coords[0]), board_size[1] - Move.SGF_COORD.index(sgf_coords[1]) - 1), player=player) def __init__(self, coords: Optional[Tuple[int, int]] = None, player: str = "B"): """Initialize a move from zero-based coordinates and player""" diff --git a/gui/badukpan.py b/gui/badukpan.py index 7957e86..29021d4 100644 --- a/gui/badukpan.py +++ b/gui/badukpan.py @@ -85,7 +85,7 @@ class BadukPanWidget(Widget): if nodes_here and max(yd, xd) < self.grid_size / 2: # load old comment if touch.is_double_tap: # navigate to move katrain.game.set_current_node(nodes_here[-1]) - self.draw_board_contents() + katrain.update_state() else: # load comments katrain.log(f"\nAnalysis:\n{nodes_here[-1].analysis}", OUTPUT_DEBUG) katrain.log(f"\nParent Analysis:\n{nodes_here[-1].parent.analysis}", OUTPUT_DEBUG) diff --git a/gui/popups.py b/gui/popups.py index 498f8aa..1bbc56a 100644 --- a/gui/popups.py +++ b/gui/popups.py @@ -163,7 +163,7 @@ class ConfigPopup(QuickConfigGui): engine_updates = updated_cat["engine"] if "visits" in engine_updates: self.katrain.engine.visits = engine_updates["visits"] - if {key for key in engine_updates if key not in {"max_visits", "max_time", "enable_ownership"}}: + if {key for key in engine_updates if key not in {"max_visits", "max_time", "enable_ownership","wide_root_noise"}}: self.katrain.log(f"Restarting Engine after {engine_updates} settings change") self.info_label.text = "Restarting engine\nplease wait." self.katrain.controls.set_status(f"Restarted Engine after {engine_updates} settings change.") @@ -176,8 +176,7 @@ class ConfigPopup(QuickConfigGui): new_engine = KataGoEngine(self.katrain, self.config["engine"]) self.katrain.engine = new_engine self.katrain.game.engines = {"B": new_engine, "W": new_engine} - if not old_proc: - self.katrain.game.analyze_all_nodes() # old engine was broken, so make sure we redo any failures + self.katrain.game.analyze_all_nodes() # old engine was possibly broken, so make sure we redo any failures self.katrain.update_state() Clock.schedule_once(restart_engine, 0) diff --git a/img/flaticon/next.png b/img/flaticon/next.png index c6bb8f6..8aab66a 100644 Binary files a/img/flaticon/next.png and b/img/flaticon/next.png differ diff --git a/img/flaticon/next5.png b/img/flaticon/next5.png index 8aab66a..3660230 100644 Binary files a/img/flaticon/next5.png and b/img/flaticon/next5.png differ diff --git a/img/flaticon/next999.png b/img/flaticon/next999.png index 49a1824..c6bb8f6 100644 Binary files a/img/flaticon/next999.png and b/img/flaticon/next999.png differ diff --git a/img/flaticon/previous.png b/img/flaticon/previous.png index 9e27db2..abb2c94 100644 Binary files a/img/flaticon/previous.png and b/img/flaticon/previous.png differ diff --git a/img/flaticon/previous5.png b/img/flaticon/previous5.png index abb2c94..8e1b633 100644 Binary files a/img/flaticon/previous5.png and b/img/flaticon/previous5.png differ diff --git a/img/flaticon/previous999.png b/img/flaticon/previous999.png index 21c20e3..9e27db2 100644 Binary files a/img/flaticon/previous999.png and b/img/flaticon/previous999.png differ diff --git a/katrain.kv b/katrain.kv index 27dbe11..56ee07e 100644 --- a/katrain.kv +++ b/katrain.kv @@ -229,7 +229,7 @@ valign: 'bottom' halign: 'left' text: '+0' - color: GREY + color: BUTTON_COLOR size: self.texture_size : @@ -271,18 +271,10 @@ id: range_label_top pos: root.right_edge - self.width-1, root.pos[1]+root.height*(1-root.marginy) - self.font_size text: 'B+' + str(int(root.y_scale)) -# GraphMarkerLabel: -# font_size: 0.1 * root.height -# pos: root.right_edge - self.width-1, root.bhalf - self.font_size + 1 -# text: 'B+' + str(int(root.y_scale/2)) GraphMarkerLabel: font_size: 0.1 * root.height pos: root.right_edge - self.width-1, root.mid - self.height/2 + 2 text: 'Jigo' -# GraphMarkerLabel: -# font_size: 0.1 * root.height -# pos: root.right_edge - self.width-1, root.whalf - 1 -# text: 'W+' + str(int(root.y_scale/2)) GraphMarkerLabel: font_size: 0.1 * root.height pos: root.right_edge - self.width-1, root.pos[1] diff --git a/katrain.py b/katrain.py index ca3ccfa..f6c55a4 100644 --- a/katrain.py +++ b/katrain.py @@ -17,7 +17,7 @@ from kivy.storage.jsonstore import JsonStore from kivy.uix.popup import Popup from core.ai import ai_move -from core.common import OUTPUT_INFO, OUTPUT_ERROR, OUTPUT_DEBUG, OUTPUT_EXTRA_DEBUG +from core.common import OUTPUT_INFO, OUTPUT_ERROR, OUTPUT_DEBUG, OUTPUT_EXTRA_DEBUG, OUTPUT_KATAGO_STDERR from core.engine import KataGoEngine from core.game import Game, IllegalMoveException, KaTrainSGF from core.sgf_parser import Move, ParseError @@ -47,7 +47,15 @@ class KaTrainGui(BoxLayout): self._keyboard.bind(on_key_down=self._on_keyboard_down) def log(self, message, level=OUTPUT_INFO): - if level == OUTPUT_ERROR: + if level == OUTPUT_KATAGO_STDERR: + if "starting" in message.lower(): + self.controls.set_status(f"KataGo engine starting...") + if message.startswith("Tuning"): + self.controls.set_status(f"KataGo is tuning settings for first startup, please wait." + message) + if "ready" in message.lower(): + self.controls.set_status(f"KataGo engine ready.") + print(f"[KG:STDERR]{message.strip()}") + elif level == OUTPUT_ERROR: self.controls.set_status(f"ERROR: {message}") print(f"ERROR: {message}") elif self.debug_level >= level: @@ -91,7 +99,7 @@ class KaTrainGui(BoxLayout): # AI and Trainer/auto-undo handlers cn = self.game.current_node auto_undo = cn.player and "undo" in self.controls.player_mode(cn.player) - if auto_undo and cn.analysis_ready and cn.parent and cn.parent.analysis_ready: + if auto_undo and cn.analysis_ready and cn.parent and cn.parent.analysis_ready and not cn.children and not self.game.ended: self.game.analyze_undo(cn, self.config("trainer")) # not via message loop if cn.analysis_ready and "ai" in self.controls.player_mode(cn.next_player).lower() and not cn.children and not self.game.ended and not (auto_undo and cn.auto_undo is None): self._do_ai_move(cn) # cn mismatch stops this if undo fired. avoid message loop here or fires repeatedly. @@ -179,7 +187,7 @@ class KaTrainGui(BoxLayout): self.fileselect_popup = Popup(title="Double Click SGF file to analyze", size_hint=(0.8, 0.8)).__self__ popup_contents = LoadSGFPopup() self.fileselect_popup.add_widget(popup_contents) - popup_contents.filesel.path = os.path.expanduser(self.config("sgf/sgf_load")) + popup_contents.filesel.path = os.path.abspath(os.path.expanduser(self.config("sgf/sgf_load"))) def readfile(files, _mouse): self.fileselect_popup.dismiss() @@ -189,6 +197,8 @@ class KaTrainGui(BoxLayout): self.log(f"Failed to load SGF. Parse Error: {e}", OUTPUT_ERROR) return self._do_new_game(move_tree=move_tree, analyze_fast=popup_contents.fast.active) + if not popup_contents.rewind.active: + self.game.redo(999) popup_contents.filesel.on_submit = readfile self.fileselect_popup.open() diff --git a/spec/KaTrain.spec b/spec/KaTrain.spec index ff42991..77aa90e 100644 --- a/spec/KaTrain.spec +++ b/spec/KaTrain.spec @@ -3,7 +3,8 @@ from kivy_deps import sdl2, glew block_cipher = None -# pyinstaller spec/katrain.spec --upx-dir my --noconfirm +# pyinstaller spec/katrain.spec --noconfirm +# --upx-dir my a = Analysis(['..\\katrain.py'], pathex=['C:\\Users\\sande\\Desktop\\katrain\\spec'],