網頁

2018年10月19日 星期五

Tesseract OCR

https://digi.bib.uni-mannheim.de/tesseract/
可以下載安裝版
雖然它只可以執行,不能開發程式,但還是先安裝,因為要使用它的 tessdata
等用完再移除吧
https://github.com/UB-Mannheim/tesseract/wiki/Windows-build
有一些安裝檔如何產生的說明,但它是利用 Linux 跨平台編譯產生的

使用 Vcpkg
開啟 PowerShell
git clone https://github.com/Microsoft/vcpkg.git vcpkg
cd vcpkg
.\bootstrap-vcpkg.bat
產生 vcpkg.exe
.\vcpkg install tesseract:x64-windows
產生 installed\x64-windows\tools\tesseract
.\vcpkg install tesseract:x64-windows-static
產生 installed\x64-windows-static\tools\tesseract
.\vcpkg install tesseract:x86-windows-static
有 include, dll, lib, 但卻是 3.05 版

使用 cmake, cppan, vs2017
原先使用之前的 cmake(3.10版), 一直失敗, 更新成 cmake(3.12版)才成功
下載 cppan
cppan 會使用 c:\users\userName\.cppan 目錄,若有失敗要重新開始,刪除這個目錄
設定 PATH 到 cmake 和 cppan
開啟 PowerShell
git clone https://github.com/tesseract-ocr/tesseract tesseract
cd tesseract
mkdir win64
cd win64
PS D:\Tesseract\tesseract\win64> $env:Path += ";D:\Tesseract\cppan-master-Windows-client;C:\Program Files\CMake\bin"
PS D:\Tesseract\tesseract\win64> $env:path.split(";")
cppan ..
cmake .. -G "Visual Studio 15 2017 Win64"
開啟 vs2017
開啟 tesseract\win64\tesseract.sln
先編譯 "CPPAN Targets/Service/cppan-d-b-d" 專案,會產生錯誤
最主要為程式內含有錯誤的字元
開啟這些檔案,另存新檔,選擇 Save 旁邊的小按鈕,選擇 Save with encoding
Encoding 選擇 Unicode (UTF-8 with signature)
ALL_BUILD 可以成功,接著 build INSTALL
此時會產生 MSB307 setlocal 錯誤
主要是因為沒有權限安裝程式到 C:\Program Files\tesseract
使用 Administrator 身分重新開啟 vs2017
重新 build 即可
增加中文字(含手寫)的支援
到 https://github.com/tesseract-ocr/tessdata 下載 tessdata
但是我不知道要下載那些檔案,乾脆使用安裝檔內的 tessdata
設定環境變數 TESSDATA_PREFIX=C:\Program Files\tesseract\tessdata



發現在部分電腦上速度會非常慢,可關閉 openmp 改善
修改 project libtesseract 和 tesseract 的 property
C/C++/Language/Open MP Support: No(/openmp-)

2018年10月6日 星期六

build darknet yolo

git clone https://github.com/AlexeyAB/darknet.git

下載 CUDA Toolkit 9.1
https://developer.nvidia.com/cuda-toolkit-archive
安裝失敗,請參考下列步驟
https://yingrenn.blogspot.com/2018/07/cuda.html

下載 cuDNN, 請選擇 cuDNN v7.0 for CUDA 9.1 for 正確 Windows 版本
https://developer.nvidia.com/rdp/cudnn-archive

使用 VS2015 開啟 D:\Tensorflow\Yolo\darknet\build\darknet\darknet.sln
切換 Win32 到 x64
Project/darknet properities/
輸入 Configuration Properties/"CUDA C/C++"/CUDA Toolkit Custom Dir
C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v9.1
輸入 Additional Include Directories
D:\Tensorflow\Yolo\opencv\build\include
D:\CUDNN\cudnn-9.1-windows7-x64-v7\cuda\include
輸入 Additional Library Directories
D:\Tensorflow\Yolo\opencv\build\x64\vc14\lib
D:\CUDNN\cudnn-9.1-windows7-x64-v7\cuda\lib\x64

2018年10月5日 星期五

git 學習紀錄

Git 只能管理單純文字檔,不能處理MS Word, pdf, 圖片或執行檔等非文字檔

安裝
# 如果你的 Linux 是 Ubuntu:
$ sudo apt-get install git-all
# 如果你的 Linux 是 Fedora:
$ sudo yum install git-all

Untracked: 在工作目錄內,尚未被追蹤或不需追蹤的檔案
Unmodified: 在工作目錄內,尚未被修改的檔案
Modified: 在工作目錄內,已經被修改的檔案
Staged: 在工作目錄內,被 add 到版本庫,尚未被 commit

設定使用者和email
$ git config --global user.name "UserName"
$ git config --global user.email "username@email.com"

在工作目錄上建立版本庫,會產生 ".git" 目錄
$ git init

查詢版本庫狀態
$ git status

將檔案放入版本庫的 stage 內
$ git add file.py
將資料夾中所有未被添加的檔案,放入版本庫的 stage 內
$ git add .

確認提交
$ git commit -m "修改說明"
確認提交(可省掉 git add)
$ git commit -am "修改說明"

查詢紀錄
$ git log
查詢紀錄,每個 commit 一行
$ git log --oneline
查詢紀錄,每個 commit 一行,並顯示 branch
$ git log --oneline --graph

查詢這次還沒add(unstaged)的修改部分 和 上個已經commit或已經add(staged)的文件差異
$ git diff
查詢已經add(staged)的修改部分 和 上個已經commit的文件差異
$ git diff --cached
查詢這次還沒add(unstaged)的修改部分 和 上個已經commit的文件差異
$ git diff HEAD

checkout 某一(id)版本,到工作目錄(HEAD 指到 id)
$ git checkout id
checkout 最新版本,到工作目錄(HEAD 指到最新 id)
$ git checkout master
checkout 某一(id)版本的 file.py
$ git checkout id -- file.py

查看所有 HEAD 改動
$ git reflog

直接修改上個 commit,不說明
$ git commit --amend --no-edit

reset 回到上一次的 commit
$ git reset --hard HEAD
reset 回到上上一次的 commit
$ git reset --hard HEAD^
reset 回到某一次 commit
$ git reset --hard id
reset 可搭配三種參數 soft, mixed(default), 以及 hard
soft 只回復 Repository
mixed 回復 staged 和 Repository
hard 回復 staged 和 Repository 和工作目錄(通常使用這個)

查詢分支
$ git branch
查詢本地和遠端的分支
$ git branch -a
查詢遠端的分支
$ git branch -r
建立分支 test
$ git branch test
checkout 分支 test
$ git checkout test
建立分支 test 並且切換(checkout)到 test
$ git checkout -b test

合併分支
先切回 master
$ git checkout master
將 test 合併至 master
$ git merge --no-ff -m "合併說明" test

合併有衝突時,先修改檔案,再 commit, 衝突就解決了
$ git commit -am "解決說明"

暫存修改
$ git stash
查詢暫存
$ git stash list
回復暫存
$ git stash pop

新增遠端節點 origin
$ git remote add origin https://github.com/UserName/git-name.git
推送本地的 master 分支到 origin
$ git push -u origin master
推送本地的 test 分支到 origin
$ git push -u origin test
本地端修改
$ git commit -am "修改說明"
推送修改到 origin
$ git push -u origin master
取回最新的遠端資料到本地
$ git pull origin master

2018年7月9日 星期一

CUDA 安裝失敗

CUDA 安裝失敗,通常是由於 Visual Studio Integration 失敗
所以透過自訂安裝,跳過不安裝 Visual Studio Integration, 可以安裝成功
Installer Type 要選擇 exe(local)

而 Visual Studio Integration 的安裝方式如下:
1. 使得可以編譯 CUDA 程式
注意安裝 CUDA 時的路徑,拷貝出 CUDAVisualStudioIntegration 目錄夾
將 D:\CUDAVisualStudioIntegration\extras\visual_studio_integration\MSBuildExtensions
目錄下所有檔案拷貝至
C:\Program Files (x86)\MSBuild\Microsoft.Cpp\v4.0\V140\BuildCustomizations
2. 使得 Visual Studio 可以新建 CUDA 專案
將目錄
D:\CUDAVisualStudioIntegration\extras\visual_studio_integration\CudaProjectVsWizards
拷貝至
C:\Program Files (x86)\Microsoft Visual Studio 14.0\Common7\IDE\Extensions
3. 安裝
D:\CUDAVisualStudioIntegration\NVIDIA_Nsight_Visual_Studio_Edition_Win64_5.4.0.17229.msi

2018年7月3日 星期二

tensorflow audio recognition 之 SpeechActivity.java

分為 record thread 和 recognize thread, 兩個 thread 依靠 recordingBuffer 交換資料
兩者速度不會一致,所以 recognize thread 可能重複 recognize, 也可能漏

short[] recordingBuffer = new short[RECORDING_LENGTH];
int recordingOffset = 0;

private void record() {
  int numberRead = record.read(audioBuffer, 0, audioBuffer.length);
  int maxLength = recordingBuffer.length;
  int newRecordingOffset = recordingOffset + numberRead;
  //int secondCopyLength = Math.max(0, newRecordingOffset - maxLength);
  if (newRecordingOffset > maxLength) {
    secondCopyLength = newRecordingOffset - maxLength;
  } else {
    secondCopyLength = 0;
  }
  int firstCopyLength = numberRead - secondCopyLength;
  System.arraycopy(audioBuffer, 0, recordingBuffer, recordingOffset, firstCopyLength);
  System.arraycopy(audioBuffer, firstCopyLength, recordingBuffer, 0, secondCopyLength);
  recordingOffset = newRecordingOffset % maxLength;
}

private void recognize() {
  int maxLength = recordingBuffer.length;
  int firstCopyLength = maxLength - recordingOffset;
  int secondCopyLength = recordingOffset;
  System.arraycopy(recordingBuffer, recordingOffset, inputBuffer, 0, firstCopyLength);
  System.arraycopy(recordingBuffer, 0, inputBuffer, firstCopyLength, secondCopyLength);
}

Yolo

目錄 data/img

檔案 data/obj.data
classes= 2
train  = data/train.txt
valid  = data/train.txt
names = data/obj.names (相對於執行檔目錄)
backup = backup/

檔案 data/obj.names
air
bird

檔案 data/train.txt
data/img/air1.jpg
data/img/air2.jpg
data/img/air3.jpg

檔案 yolo-obj.cfg
(測試用)
batch=1
subdivisions=1
(訓練用)
batch=64
subdivisions=1, (視記憶體大小修改,記憶體小則使用64)
修改所有 [yolo] 層內的
classes =
修改所有 [yolo] 前一個 [convolutional] 層內的
filters = (classes + 5) * 3

標記
yolo_mark.exe data/img data/train.txt data/obj.names

訓練
darknet.exe detector train data/obj.data yolo-obj.cfg darknet19_448.conv.23
obj.data 內的 backup 指定輸出 weights 存放位置
darknet19_448.conv.23: 其實就是 weights, 要接續中斷的訓練時,則改為新產生的 weights
-dont_show: 不顯示 Loss-Window

檢測訓練結果(IoU, mAP)
darknet.exe detector map data/obj.data yolo-obj.cfg backup\yolo-objj_7000.weights

COCO Yolo v3(4GB GPU): yolov3.cfg, yolov3.weights
COCO Yolo v3 tiny(1GB GPU): yolov3-tiny.cfg, yo.ov3-tiny.weights
COCO Yolo v2(4GB GPU): yolov2.cfg, yolov2.weights
VOC Yolo v2(4GB GPU): yolo-voc.cfg, yolo-voc.weights
COCO Yolo v2 tiny(1GB GPU): yolov2-tiny.cfg, yolov2-tiny.weights
VOC Yolo v2 tiny(1GB GPU): yolov2-tiny-voc.cfg, yolov2-tiny-voc.weights
以上似乎是訓練時的需求,檢測或分類時似乎沒那麼大的需求

darknet.exe 參數
-i <index>, 指定 GPU, 可用 nvidia-smi.exe 查詢
-nogpu, 不使用 GPU
-thresh <val>, 預設為 0.25
-c <num>, OpenCV 影像, 預設為 0
-ext_output, 輸出物件位置
detector test, 相片
detector demo, 影片
detector train, 訓練
detector map, 檢測訓練結果
classifier predict, 分類

./darknet detect cfg/yolov3.cfg yolov3.weights data/dog.jpg
./darknet detector test cfg/coco.data cfg/yolov3.cfg yolov3.weights data/dog.jpg
以上兩個命令一樣

使用命令取得 yolov3-tiny.conv.15
darknet.exe partial cfg/yolov3-tiny.cfg yolov3-tiny.weights yolov3-tiny.conv.15 15

如何增進物件檢測
訓練前:
.cfg 檔內的 random=1
增加 .cfg 檔內的 width, height (須為 32 的倍數)
執行下列命令,重新計算 anchors, 更改 .cfg 檔內的 anchors
darknet.exe detector calc_anchors voc.data -num_of_clusters 9 -width 416 -height 416
小心標註相片內的物件,每一物件都要標註,而且不要標錯
每個物件最好有 2000 以上的影像,包含有不同的大小、角度、光線、背景等
不要被檢出的物件要在相片內,而且不能被標註

訓練時相片和標註檔的對映
darknet.c
int main(int argc, char **argv)
>run_detector(argc, argv);
detector.c
void run_detector(int argc, char **argv)
>train_detector(datacfg, cfg, weights, gpus, ngpus, clear, dont_show);
void train_detector(char *datacfg, char *cfgfile, char *weightfile, int *gpus, int ngpus, int clear, int dont_show)
> pthread_t load_thread = load_data(args);
data.c
pthread_t load_data(load_args args)
>if(pthread_create(&thread, 0, load_threads, ptr)) error("Thread creation failed");
void *load_threads(void *ptr)
>threads[i] = load_data_in_thread(args);
if(pthread_create(&thread, 0, load_thread, ptr)) error("Thread creation failed");
void *load_thread(void *ptr)
>*a.d = load_data_detection(a.n, a.paths, a.m, a.w, a.h, a.c, a.num_boxes, a.classes, a.flip, a.jitter, a.hue, a.saturation, a.exposure, a.small_object);
data load_data_detection(int n, char **paths, int m, int w, int h, int c, int boxes, int classes, int use_flip, float jitter, float hue, float saturation, float exposure, int small_object)
>fill_truth_detection(filename, boxes, d.y.vals[i], classes, flip, dx, dy, 1./sx, 1./sy, small_object, w, h);
void fill_truth_detection(char *path, int num_boxes, float *truth, int classes, int flip, float dx, float dy, float sx, float sy, int small_object, int net_w, int net_h)
>replace_image_to_label(path, labelpath);
utils.c
void replace_image_to_label(char *input_path, char *output_path)

在相片上標註偵測出的物件
image.c
void draw_detections_cv_v3(IplImage* show_img, detection *dets, int num, float thresh, char **names, image **alphabet, int classes, int ext_output)

network.c
將 image 轉成 network
float *network_predict(network net, float *input)
從 network 中取得 detection
detection *get_network_boxes(network *net, int w, int h, float thresh, float hier, int *map, int relative, int *num, int letter)



2018年6月8日 星期五

Training your Object Detection Classifier

Download TensorFlow Models
Download object detection models from model zoo

set PATH=%PATH%;..\..\protoc-3.5.1-win32\bin
cd models-master\research
protoc.exe object_detection\protos\anchor_generator.proto --python_out=.
protoc.exe object_detection\protos\argmax_matcher.proto --python_out=.
.
.
.
protoc.exe object_detection\protos\train.proto --python_out=.
cd ..\..\

cd models-master\research
python setup.py build
python setup.py install

修改 models-master/research/object_detection/trainer.py
    # Soft placement allows placing on CPU ops without GPU implementation.
    session_config = tf.ConfigProto(allow_soft_placement=True,
                                    log_device_placement=False)
    session_config.gpu_options.allow_growth=True
    session_config.gpu_options.allocator_type = "BFC"
    session_config.gpu_options.per_process_gpu_memory_fraction = 0.4


LabelImg 標註相片,產生 .xml
修改 xml_to_csv.py 由 .xml 產生 .csv
修改 generate_tfrecord.py 內的 class_text_to_int(row_label), 標註從 1 開始,不是 0
由 .csv 產生 .record(TFRecord 檔)
準備 labelmap.pbtxt, 標註從 1 開始,不是 0

準備修改 models-master\research\object_detection\samples\configs