ラベル 遊び の投稿を表示しています。 すべての投稿を表示
ラベル 遊び の投稿を表示しています。 すべての投稿を表示

2023年9月28日木曜日

ChatGPTを研究に利用するための検討 #1

article_230926

どのように研究や解析に利用できるか

今年度、専門家だけでなく、広く一般に話題となっている科学技術のトピックスとしては、OpenAIなどが進める生成系AIが挙げられる。

生成AI(または生成系AI)とは、「Generative AI:ジェネレーティブAI」とも呼ばれ、さまざまなコンテンツを生成できるAIのことです。 従来のAIが決められた行為の自動化が目的であるのに対し、生成AIはデータのパターンや関係を学習し、新しいコンテンツを生成することを目的としています。 https://www.nri.com/jp/knowledge/glossary/lst/sa/generative_ai

引用したリンク先の解説でも述べられていたが、劇的な精度向上と一般的なユーザーでも導入して利用可能になってきたことで、社会に大きな影響を与えつつあると思う。では、このような生成系AIの技術は、研究や解析の活動にどれくらい使用できるものなのだろうか?

周りの人と話をしていると、すでにコードを書いたり何か研究に関係することを調べたりするのに使用しているという話を意外と聞くようになった。スクリプトを描くときにたたき台を作るのに便利という話も聞き、興味が湧いてきた。

すでに先行して利用されている方々の解説や記事はたくさんあると思うので、遅れての検討となるのだろうが、バイオ系の研究に関連して、どのような利用方法があるのかを検討してみたいと思い、ChatGPTを使用し始めた。

ChatGPTへの登録

まずは、ChatGPTのサイトからログインのためのアカウントを作成した。以下のサイトから、手順に乗って手続きした。

生成系のAIは現在注目の的になっていて、さまざまなサイト、アプリが似たような名前でサービスを展開しているようだ。プログラミングなど、より特定の目的に合致したサービスがあるかもしれないが、とりあえず、有名になったOpenAIのChatGPTを使用してみることにした。

Python script assistance

聞きかじった情報を基に、以下のようなプロンプト文を作成してみた。このようなプロンプトにより方向性を作った方がより関連性の深い内容を生成してくれるらしく、とりあえずそれらしいプロンプトを考えた。

Help python scripting: You are good skilled python programmer. You are also good instructor of python programming. Your mission is to help me programming python script.

Keep in mind of the following items. - If there are mis programming in the input, debug them using standard python grammer. - Simple, easy-to-read scripting. - When you debug or revise scripts, explain the reason.

Of course, I’d be happy to help you with Python scripting! Please provide the specific script or programming task you need assistance with, and I’ll do my best to guide you through it, explain any issues, and provide clear, easy-to-read code. Just share the code or describe the problem you’re facing, and we can get started.

何か適当な例はないかと考え、pythonでk-meansクラスタリングを実装する方法を解説してもらう、というお題を考えた。
まず、以下のような質問を行った。

I have a problem with python scripting. How can I do K-means clustering for data of data frame with x and y columns?

To perform K-means clustering on a DataFrame with ‘x’ and ‘y’ columns in Python, you can use the KMeans algorithm provided by the popular library, scikit-learn. Here’s a step-by-step guide on how to do this:

  1. Import the necessary libraries:
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
  1. Load your data into a DataFrame. Let’s assume you have a CSV file named ‘data.csv’ with ‘x’ and ‘y’ columns.
data = pd.read_csv('data.csv')
  1. Select the ‘x’ and ‘y’ columns as your feature set:
X = data[['x', 'y']]
  1. Choose the number of clusters (K) you want to create. This is a hyperparameter that you need to decide based on your specific problem or by using techniques like the elbow method.
k = 3  # You can adjust this value based on your needs
  1. Create a K-means model and fit it to your data:
kmeans = KMeans(n_clusters=k)
kmeans.fit(X)
  1. Assign cluster labels to each data point:
data['cluster'] = kmeans.labels_
  1. You can now analyze the results. For example, you can print the centroids of each cluster:
centroids = kmeans.cluster_centers_
print("Cluster centroids:")
for i, centroid in enumerate(centroids):
    print(f"Cluster {i+1}: {centroid}")
  1. Visualize the clustered data using matplotlib:
plt.scatter(data['x'], data['y'], c=data['cluster'], cmap='rainbow')
plt.scatter(centroids[:, 0], centroids[:, 1], c='black', marker='x', s=100)
plt.xlabel('X')
plt.ylabel('Y')
plt.title('K-means Clustering')
plt.show()

This script will load your data, perform K-means clustering with the specified number of clusters, and visualize the results. You can adjust the k value to change the number of clusters according to your specific data and requirements.

If you encounter any errors or have further questions, please provide more details about your problem, and I’ll be happy to assist you further.

一気にスクリプト例までアウトプットしてくれている。これに対して、実装を試してみるためのサンプルデータが必要だと思うが、その方法についても尋ねてみた。以下のような質問を行った。

Can you tell me how to generate sample data?

Certainly! You can generate sample data for your K-means clustering experiment in Python using libraries like NumPy and pandas. Here’s how you can create synthetic data with random values for ‘x’ and ‘y’ columns:

import numpy as np
import pandas as pd

# Set the random seed for reproducibility
np.random.seed(42)

# Define the number of data points you want
num_points = 100

# Generate random 'x' and 'y' values
x_values = np.random.uniform(0, 100, num_points)  # Replace 0 and 100 with your desired range
y_values = np.random.uniform(0, 100, num_points)  # Replace 0 and 100 with your desired range

# Create a DataFrame
data = pd.DataFrame({'x': x_values, 'y': y_values})

# Display the first few rows of the generated data
print(data.head())

# Save the generated data to a CSV file if needed
data.to_csv('sample_data.csv', index=False)

In this example:

  • We use np.random.uniform to generate random ‘x’ and ‘y’ values within the specified range (0 to 100 in this case). You can adjust these ranges according to your data distribution.

  • We create a DataFrame using pd.DataFrame and provide the ‘x’ and ‘y’ values as columns.

  • If you want to save this generated data to a CSV file, you can use to_csv() as shown.

Now you have sample data with random ‘x’ and ‘y’ values that you can use for your K-means clustering experiments. You can adjust the number of data points, the range of values, and other parameters to suit your specific needs.

結構懇切丁寧に解説してくれるのだなと思った。また、出力がコードブロックを使って読みやすく示され、結果をコピーしてすぐに使用できるようになっていることも便利に思えた。

ChatGPT Promptについて

これまで、プロンプトを特に決まり等を調査せずに作成していたが、実際にはもっと良い書き方があるのかもしれない。
検索して様々な活用事例を見ているのだが、ChatGPTそのものにも聞いてみようと思い、新しいchatを立ち上げて、以下のプロンプトを作成した。

You are skilled IT engineer good at using ChatGPT. Help me with my question about using ChatGPT service.

Of course, I’d be happy to help you with your question about using ChatGPT! Please go ahead and ask your question, and I’ll do my best to provide you with the information and guidance you need.

I am writing chatGPT prompt. Please check the following chatGPT prompt and, if necessary, revise the following chatGPT prompt.

I’d be happy to help you review and revise your ChatGPT prompt. Please provide the prompt you’d like me to check, and I’ll give you feedback and suggestions to make it more effective or clarify any issues.

ここで、最初に使用したプロンプト文を入力して、ChatGPTに書き直させてみた。

Help python scripting: You are good skilled python programmer. You are also good instructor of python programming. Your mission is to help me programming python script.

Keep in mind of the following items. - If there are mis programming in the input, debug them using standard python grammar. - Simple, easy-to-read scripting. - When you debug or revise scripts, explain the reason.

Your prompt is generally clear, but it could be more specific to get the best results from ChatGPT. Here’s a slightly revised version of your prompt:

“I need help with Python scripting. I’m looking for a skilled Python programmer who can also provide clear explanations. Please assist me with writing and debugging Python scripts.

Here are some important considerations: 1. If there are any programming errors in the input, please debug them using standard Python syntax. 2. I prefer simple and easy-to-read scripting. 3. Whenever you debug or revise scripts, please explain the reason for the changes.”

This revised prompt provides a clear request for assistance with Python scripting and highlights the importance of clear explanations and clean, readable code. It should help you get more focused and helpful responses from ChatGPT.

リバイスされた文章を読む限り、英語のこなれてなさが消え、より読みやすくなっているように感じられた。また、箇条書きの部分は数字に変化していた。

chatGPT promptのフォーマット、style guideについて

そもそも、このプロンプト文の作成には何か特定のフォーマットやスタイルガイドがあるのだろうか?その辺りあまり調べずに始めているのだが、それ自体も質問してみようと思い、以下の文章を入力した。

Thank you. I will try. By the way, is there any format or style guide of chatGPT prompt?

As of my last knowledge update in September 2021, there wasn’t an official format or style guide specifically provided by OpenAI for ChatGPT prompts. However, there are some general principles you can follow to create effective prompts:

  1. Be Clear and Specific: Clearly state what you want from ChatGPT. Specific questions or instructions tend to yield more precise responses.

  2. Provide Context: If your request involves a specific topic or domain, provide context. ChatGPT can generate better responses when it understands the context of your question.

  3. Be Concise: Avoid overly long and complex prompts. Keep your requests concise and to the point.

  4. Ask Open-Ended Questions: If you want a detailed response, ask open-ended questions that require more than a yes/no answer.

  5. Request Explanations: If you want detailed explanations or reasoning, explicitly ask for them in your prompt.

  6. Politeness: Being polite and using respectful language can lead to more cooperative responses.

  7. Experiment: Don’t be afraid to experiment with different prompts to see what works best for your specific needs.

Please note that the effectiveness of your prompt can also depend on the model’s capabilities and how it has been fine-tuned, so it’s a good idea to iterate and refine your prompts based on the responses you receive.

OpenAI might have released updated guidelines or recommendations since my last update, so it’s a good practice to check their official documentation or guidelines for any new information regarding prompt formatting and style.

少なくとも21年の段階では、決まったフォーマットやスタイルがあるわけではないらしいが、色々と試行錯誤が必要らしい。また、公式のドキュメントがあるのかについては確認してみたい。少し面白かったのは、6番目のplitenessだ。AIに対して尋ねるのに、この項目があるのはよく考えてみると少し奇妙な気がする。ただ、見返してみるとこの回答を見る前の段階で、自分の入力分には”please”や”thank you”といった人と話している時と同じような書き方がしてあって、自分自身不思議に思った。より無礼な聞き方をすると結果が変わるのだろうか?今後試してみたいような気もする。

2021年5月7日金曜日

References of Ubuntu 20.04 Installation

Description 

Ubuntuマシンを使いたくなり、久しぶりにLinuxのクリーンインストールを実施した。手法の詳細は様々な情報があるので、差し当たり自分が使用した情報のreferenceを覚書として記録した。


References

ubuntu.jp

  1. https://www.ubuntulinux.jp

ブートUSB作成(Mac)

  1. https://sy-base.com/myrobotics/mac/ubuntu_bootusb/

ubuntu 日本語Remixダウンロードサイト

  1. https://www.ubuntulinux.jp/download

インストール情報

  1. https://linuxfan.info/ubuntu-20-04-install-guide#google_vignette

日本語入力変換

  1. https://linuxfan.info/ubuntu-18-04-japanese-input

解像度問題〜NVIDIA gpu  driver更新

  1. https://ayatansnl.hatenablog.com/entry/2016/03/02/190055
  2. https://zenn.dev/190ikp/articles/how_to_install_nvidia_drivers
  3. https://qiita.com/kenji-miyake/items/06b8c3807bef0ba5c451#nvidia-driver-のインストール
  4. https://hibiki-press.tech/dev-env/ubuntu/add-ppa/4640
  5. https://qiita.com/ksasaki/items/b20a785e1a0f610efa08
  6. https://k-hyoda.hatenablog.com/entry/2020/07/09/223907
  7. https://linuxconfig.org/how-to-install-the-nvidia-drivers-on-ubuntu-20-04-focal-fossa-linux

IPアドレス確認

  1. https://news.mynavi.jp/article/20190825-883094/

SSHインストールと接続

  1. https://codechacha.com/ja/ubuntu-install-openssh/
HDD, SSDマウント確認
  1. アプリケーション:「ディスク」で確認できる。マウントポイントも見れる。ここでマウントポイントの書き換えも可能らしい
  2. コマンドラインで見るときはfdiskで確認できる
  3. sudo disk -l
  4. ファイルシステムの確認は、lsblk -f
  5. マウントポイントは/media/<my_dir>/<UUID>になっていた。UUIDが長いので、この「ディスク」アプリケーションを使って書き換えることもできそう。

解析用プログラム、アプリケーション準備
  1. Atom
  2. Anaconda3 (Jupyter, Spyder)
  3. R
  4. RStudio
  5. R package各種






2019年12月10日火曜日

qiita advent calendarにバイオインフォマティクスやメタゲノミクスのカレンダー

プログラミングやコンピュータ関係のtipsを共有するためのプラットフォームサイトでQiitaというのがあり、よく見ている。

https://qiita.com
Qiitaは、プログラマのための技術情報共有サービスです。 プログラミングに関するTips、 ノウハウ、メモを簡単に記録 & 公開することができます。

単純に困った時に調べるために使うことも多いが、ぼーっと見て回って面白い記事を見つけて楽しむことも多い。Qiitaは12月にクリスマスまでの日数をカウントダウンするアドベントカレンダーのサイトを設置している。

https://qiita.com/advent-calendar/2019
Qiita Advent Calendarとは、クリスマスまでの日数をカウントダウンするアドベントカレンダーの習慣にもとづいて毎年12月1日から25日までの期間限定で展開される記事投稿イベントです。毎年、Qiitaとクリスマスを最高に盛り上げる一大イベントとなっております。興味のあるトピックのカレンダーに参加し、この年末を最高に盛り上がる年末にしていきましょう🎉豪華景品がもらえるスポンサーカレンダーもありますので、ぜひ奮ってご参加ください🎄

12/1から12/25まで参加者がトピックのカレンダーを作って、記事をアップしていくというもので、プログラミング言語やデータベース、アプリケーションなどのガチIT系のカレンダーに加えて、学術というカテゴリーのカレンダーも30ほど作られている。今見てみると300人を超える参加者がいるようだ。学術にカテゴライズされるカレンダーにはDeep Learningや機械学習などのコンピュータサイエンス系のものに混じって、メタゲノムやバイオインフォマティクスなど、生物系のトピックスも立っており面白い。

例えば、メタゲノムの12/3の記事は「なぜメタゲノム解析するのか?」

https://qiita.com/i-za-world/items/6b448b52703a23e14a1b
メタゲノマーによるメタゲノマーの為のアドベントカレンダーではありますが、アドベントカレンダー初日なので、ここにいざわーるど解釈の「メタゲノム解析って何?」的なポエム記事を書いておきます。
バイオインフォマティクスのカレンダーは人気らしく、既に全ての日が埋まっているみたいだ。

https://qiita.com/advent-calendar/2019/bioinfo
バイオインフォマティクス(生命情報科学)分野のアドベントカレンダーです。分野の定義など細かいことは問いません。アルゴリズム、ツールの使い方、データベース関係、可視化、計算機の運用、バイオインフォに役立つライブラリやプログラミングテクニック、ワークフローエンジンなどなど、あなたなりのバイオインフォマティクスの知見を共有してください。


12/9の記事は「細菌のゲノムを一つだけ渡された時、どうするか」

https://qiita.com/xvtyzn/items/49cf51b6e6fee9df41af

細菌ゲノムの解析に使うツールなどの説明がされている。細菌ゲノムとかメタゲノムは触ったことがないので知らないツールが多い。色々解説してもらえているので、一通り見てみたいなと思った。

12/23にはナノポアの解析ツールの記事が予定されているようで、これもぜひ読んでみたい。

全然生物とは関係ないが、天文学というカレンダーがある。

https://qiita.com/advent-calendar/2019/tenmongaku

12/3の記事に「赤外線天文衛星「あかり」などの望遠鏡によって撮影された写真をいじってみたり」があり、なんか面白そう。写真も取得できるらしく、丁寧にスクリプトが書かれているので、時間ある時やってみたい。なかなか時間ないけど。